تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
Pydantic AILLMsFrameworks & agents3 مستويات101 قسمًايغطّي Pydantic AI 2.52دليل بالإنجليزية

The Complete Pydantic AI Guide

Build type-safe agents and structured-output LLM apps in Python with Pydantic AI. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
19sections
36examples

This is part one of three. It covers everything you need to build a real, working AI agent with Pydantic AI, not a teaser. By the end you can call a language model from typed Python, force it to answer in a structure your code can trust, give it tools it can call, pass it your database connection safely, hold a multi-turn conversation, stream its answer, and test all of it without spending a cent. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and an agent only starts to make sense once you have watched your own one call a tool, fail validation, and recover.

One warning before we begin. Pydantic AI released its second major version in June 2026, and a great deal of the code on blogs and in older tutorials was written for version 1. This guide targets Pydantic AI 2.52. Where an old habit now breaks, the text says so and shows the current form, and there is a short section near the end on recognising version 1 code.

What Pydantic AI is, and the problem it solves

A language model, on its own, takes text in and returns text out. That is enough for a chat window and not nearly enough for software. Software wants a number it can add up, a record it can store, a decision it can branch on. Software also wants the model to look things up, call a service, or check a database, and to do so without a developer hand-writing the same fragile loop for the hundredth time.

Pydantic AI is a Python library for building agents on top of language models. An agent, in this library, is a small object that sends your instructions and a question to a model, lets the model call Python functions you wrote, checks everything that comes back, and keeps going until it has an answer in exactly the shape you asked for. The whole thing is typed, so your editor and your type checker understand what goes in and what comes out.

The library comes from the team behind Pydantic, the validation library that sits underneath a large part of the Python ecosystem, and that heritage is the point. Pydantic turns "a dictionary that is probably what I expected" into "an object I have verified". Pydantic AI applies the same idea to the least predictable component you will ever put in a program, a language model. The model proposes; Pydantic validates; if the proposal is wrong, the validation error goes back to the model and it tries again.

What came before? In the early days you called a vendor's HTTP API directly, built a prompt string by hand, and parsed the reply with regular expressions and hope. Then came larger frameworks that offered chains, graphs, memory modules and dozens of integrations. They are powerful, but newcomers often find that a simple task needs a lot of vocabulary. Pydantic AI takes a narrower line: ordinary Python, ordinary type hints, ordinary functions as tools, and a small number of concepts. If you have used FastAPI, the feeling is similar, and the team says so openly; the documentation compares agents to a FastAPI app, something you create once and reuse.

YOUR QUESTIONplus instructions
→
MODELproposes text or a tool call
→
PYDANTICvalidates the proposal
→
TYPED RESULTor a retry

Three things about this design explain most of what follows.

The model is untrusted input. Whatever a model returns is just text that looks plausible. Pydantic AI treats it the way a web framework treats a form submission: validate it, and reject it if it does not fit. This is why the library feels strict compared with simply calling a model and printing the result.

Everything is driven by types. The shape of the answer is a Python class. The arguments of a tool are the parameters of a Python function. The data you hand to your tools is a dataclass. The library reads those types and builds the schemas it sends to the model, so you write each fact once.

It is a library, not a service. There is no server to run, no daemon, no database to install. An agent lives inside your own program, whether that is a script, a notebook, a FastAPI endpoint or a background worker. It calls a model provider over the network and that is the only external moving part.

You need little to follow along: Python 3.10 or newer, a terminal, and an account with one model provider. If you have no key yet, do not worry. A built-in test model lets you do most of this guide with no key and no network at all.

Try it
  1. Think of one task where you now paste text into a chatbot and copy the answer back by hand.
  2. Write down the exact fields you would want if a program received the answer instead (names, numbers, a yes or no).
  3. Note one thing the model would need to look up that it cannot know on its own.
a small record definition plus a lookup. That pairing, a typed result and a tool, is precisely what Pydantic AI automates.

The mental model: five nouns

Pydantic AI has a larger surface than this guide, but you can do real work by understanding five ideas. Learn these and the rest of the documentation becomes easy to navigate.

Agent. The Agent class is the central object. It holds a default model, your instructions, the tools the model may use, the type of result you expect, and the type of the dependencies you will pass in. It is generic: the full type is Agent[DepsT, OutputT], which reads "an agent that takes dependencies of this type and produces output of that type". You create it once, usually at module level, and call it many times. It keeps no memory between calls on its own.

Model. The model is the language model that does the thinking, and the library speaks to many of them: OpenAI, Anthropic, Google Gemini, Amazon Bedrock, Groq, Mistral and others, including models on your own machine through Ollama. You usually name one with a short string of the form provider:model-name, such as openai:gpt-5.2. Behind that string sits a class that knows one vendor's API.

Tool. A tool is a Python function the model is allowed to call. You write the function and decorate it; the library describes it to the model from its signature and its docstring. When the model decides it needs the function, it names it and supplies arguments, the library validates those arguments, runs your function, and sends the return value back to the model.

Output. The output is the final, typed answer. By default it is a plain string. You can instead ask for a number, a list, or any Pydantic model, dataclass or typed dictionary, and the library makes sure the model's answer fits before you ever see it.

Dependencies. Dependencies, or deps, are the things your tools need from the outside world: a database connection, an HTTP client, an API key, the identity of the current user. You pass them in for each call and your tools read them from a context object. This keeps tools testable and keeps secrets out of global variables.

Around those five sits a sixth term you will meet constantly, the run. A run is one execution of the agent: you give it a prompt, it loops through model calls and tool calls, and it ends with a result or an error. A conversation is simply several runs where you pass the earlier messages back in.

Here is the loop in more detail, because understanding it makes every error message easier to read.

1. SENDinstructions, history, prompt
→
2. MODEL REPLIEStext or tool calls
→
3. RUN TOOLSvalidated arguments
→
4. REPEATuntil valid output

First the agent sends the model your instructions, any earlier messages, and the new prompt. The model answers either with text or with a request to call one or more tools. If it asked for tools, the library checks the arguments against each tool's schema. If they are valid, it runs the functions and sends the results back as a new message, then asks the model again. If the arguments are invalid, it sends the validation error back instead and asks again. This repeats until the model produces something that satisfies the output type, or a limit is reached.

A final detail that surprises people: when you ask for structured output, the library normally delivers the output schema to the model as a tool as well. The model "calls" a tool whose arguments are your result fields, and validation of that call is what produces your typed object. You never need to think about this when things work, but it explains why output validation errors look like tool errors.

Try it
  1. Take the task from the previous section and label each part with one of the five nouns.
  2. Which part is the output type? Which part is a tool? What would you pass as dependencies?
a clean split: the shape of the answer, one or two functions the model may call, and the handles those functions need.

Installing and checking your setup

Pydantic AI needs Python 3.10 or newer and installs with pip or uv. Both work on Linux, macOS and Windows, because the package is pure Python. A virtual environment keeps the library and its many dependencies away from your system Python, so use one.

On Linux or macOS:

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install pydantic-ai

On Windows, in PowerShell:

POWERSHELL
py -m venv .venv
.venv\Scripts\Activate.ps1
py -m pip install pydantic-ai

If you prefer uv, the fast Python package manager, the equivalent is:

BASH
uv init
uv add pydantic-ai

The plain pydantic-ai package is the convenient bundle. As of version 2 it installs the core library plus support for OpenAI, Anthropic and Google models, the command-line tool, MCP support (a protocol for plugging in external tools), the evaluation library, a small web interface, retry helpers and Logfire, the observability service from the same company. It does not include the extras for Amazon Bedrock, Groq, Mistral, Cohere, xAI and Hugging Face. If you want one of those, name it in square brackets:

BASH
pip install "pydantic-ai[bedrock,groq]"

There is also a slimmer package, pydantic-ai-slim, where you choose exactly what you want. It is useful in production images where size matters:

BASH
pip install "pydantic-ai-slim[openai]"

Because the project releases a new minor version every day or two, it is good practice to see what you got. Check the version from Python:

BASH
python -c "import pydantic_ai; print(pydantic_ai.__version__)"

You should see 2.52.0 or something newer in the 2.x line. Then run the smallest possible agent. It uses a built-in fake model named test, so it needs no API key and makes no network call:

smoke_test.py
from pydantic_ai import Agent

agent = Agent('test')
result = agent.run_sync('hello')
print(result.output)

If that prints a short placeholder string, your installation works. The test model is real documented behaviour, not a hack: it exists so that you and your test suite can exercise an agent without a provider.

Windows and the local workspace Everything in this guide runs on Windows. One advanced feature does not: the documentation states that the local workspace backend does not run on Windows, and tells you to use a sandbox capability or WSL (Windows Subsystem for Linux) there. You will not meet it as a beginner, but remember it if you later give an agent file or shell access.

If you work in Jupyter, Google Colab or a similar notebook, there is one difference. A notebook already runs an event loop, so calling run_sync can fail with RuntimeError: This event loop is already running. In a notebook, write result = await agent.run('hello') at the top level instead. We explain the difference between those two calls shortly.

Try it
  1. Create a fresh virtual environment and install pydantic-ai.
  2. Print the version and run the Agent('test') smoke test.
  3. Run clai --version as well, which confirms the command-line tool arrived with it.
a 2.x version number and a placeholder reply, all without any account or key.

Your first real agent

Now to talk to an actual model. You need a key from one provider, and the library reads it from an environment variable named for that provider. For OpenAI it is OPENAI_API_KEY; for Anthropic, ANTHROPIC_API_KEY; for the Gemini API, GOOGLE_API_KEY. Set it in your shell rather than writing it into code, because anything in code ends up in version control sooner or later.

BASH
export OPENAI_API_KEY='paste-your-key-here'

On Windows PowerShell the equivalent is $env:OPENAI_API_KEY="paste-your-key-here". Now the agent:

first_agent.py
from pydantic_ai import Agent

agent = Agent(
    'openai:gpt-5.2',
    instructions='Be concise.',
)

result = agent.run_sync('What is the capital of Italy?')
print(result.output)

Run it with python first_agent.py. The answer is a short sentence naming Rome. Take the program apart, because every later example is a variation on it.

The first argument, 'openai:gpt-5.2', picks the model. The part before the colon is the provider, and the part after is the model's own name. Since version 2 the provider prefix is required. Writing just 'gpt-5.2', which worked in version 1, now fails with a UserError saying the model is unknown.

The instructions argument tells the model how to behave on every run. We will say more about it shortly.

run_sync runs the whole loop and waits for it to finish. It returns an AgentRunResult, and the final answer is in its .output attribute. That result also carries the token usage and the full message exchange, which we use later.

There are in fact five ways to run an agent. As a beginner you need only the first three, but it helps to know the family exists.

Call What it is for
agent.run_sync(...) Scripts, tests, and anywhere you are not already in async code. Blocks until done
await agent.run(...) Async code: web servers, notebooks, anything with an event loop
async with agent.run_stream(...) Showing the answer as it is generated
agent.run_stream_events(...) Streaming every event, including tool calls (mid-level)
agent.iter(...) Stepping through the loop node by node (mid-level)

run_sync is a convenience wrapper around run. Use run_sync for scripts and learning, and run inside any async def.

Now look at what came back besides the text.

inspect_result.py
from pydantic_ai import Agent

agent = Agent('openai:gpt-5.2', instructions='Be concise.')
result = agent.run_sync('What is the capital of Italy?')

print(result.output)
print(result.usage)
print(len(result.all_messages()))

result.usage reports how many requests were made and how many input and output tokens were used, and it is where cost accounting starts. Note that in version 2 it is a property, so there are no parentheses. Version 1 wrote result.usage(), and copying that into 2.x raises an error. result.all_messages() returns the full list of messages exchanged, a request and a response in the simplest case, which is the raw material for conversations.

Try it
  1. Run the first agent with a real key and change the question three times.
  2. Print result.usage each time and compare the input and output token counts.
  3. Change instructions to "Answer in one word" and run the same question again.
the same agent object giving different answers for different prompts, with token counts rising with longer questions, and one changed instruction visibly changing the style.

Choosing a model and providing credentials

The string provider:model is the everyday way to pick a model, so it is worth knowing the common prefixes. In version 2 they are openai, anthropic, google for the Gemini API, google-cloud for Gemini through Google Cloud, bedrock, groq, mistral, xai, openrouter and ollama, among others. clai --list-models, which we meet later, prints the names the library knows about.

Two renames from version 1 catch people out. The old google-gla: and google-vertex: prefixes are now google: and google-cloud:. And the bare openai: prefix now means OpenAI's newer Responses API. If you specifically want the older Chat Completions API, which many OpenAI-compatible servers implement, write openai-chat:. For most beginners the default does exactly what you want, but when you point the library at a third-party server that copies the older API, openai-chat: is the one to try.

Each provider reads its credentials from environment variables. The most common are:

Provider Variable
OpenAI OPENAI_API_KEY
Anthropic ANTHROPIC_API_KEY
Google Gemini API GOOGLE_API_KEY
Groq GROQ_API_KEY
Mistral MISTRAL_API_KEY
OpenRouter OPENROUTER_API_KEY
Amazon Bedrock AWS_BEARER_TOKEN_BEDROCK, or the usual AWS access key pair

To pass a key explicitly, or to point at a different endpoint such as an OpenAI-compatible server in your own network, build the provider yourself.

custom_provider.py
from pydantic_ai import Agent
from pydantic_ai.models.openai import OpenAIChatModel
from pydantic_ai.providers.openai import OpenAIProvider

model = OpenAIChatModel(
    'my-model-name',
    provider=OpenAIProvider(
        base_url='http://localhost:8000/v1',
        api_key='not-needed-locally',
    ),
)
agent = Agent(model)

That pattern is how you use a self-hosted model, which matters for teams in regulated sectors who cannot send prompts outside their own network or country. The library does not care whether the endpoint is a cloud service or a server in your own data centre, as long as it speaks a supported API. If you run models locally with Ollama, there is a dedicated ollama: prefix, and the neighbouring guide on Ollama shows how to serve a model on your own machine. If you want one gateway in front of many providers, the guide on LiteLLM covers that idea.

What happens when you forget the key? The library tells you plainly. For OpenAI the error is a UserError that begins "Set the OPENAI_API_KEY environment variable or pass it via OpenAIProvider(api_key=...) to use the OpenAI provider." It then adds a hint: "To try Pydantic AI without an API key, use the built-in test model: Agent('test')." That is the library reminding you of the smoke test from the installation section. The same pattern holds for the other providers, naming their own variable.

Never paste a key into your code A key in a Python file ends up in git history, in screenshots, and in shared notebooks. Keep it in an environment variable or a secrets manager, add any .env file to .gitignore, and rotate the key immediately if it ever leaks.
Try it
  1. Unset your key (unset OPENAI_API_KEY on Linux or macOS) and run the first agent.
  2. Read the error text and find the sentence that names the variable to set.
  3. Set the key again and confirm it works.
a clear UserError that names exactly what is missing, which is how most Pydantic AI errors are designed to read.

Instructions: telling the agent how to behave

Instructions are the standing orders you give the model: its role, its tone, what it should and should not do. Pydantic AI offers two related features, instructions and system prompts, and the documentation recommends instructions for almost everything.

The difference is subtle but practical. Instructions are re-evaluated on every run, and only the current agent's instructions are sent. If you pass earlier messages in as history, instructions stored inside that history are not sent again. A system prompt, by contrast, is kept in the message history and travels with it. Beginners should reach for instructions, because it means you can change an agent's behaviour by editing code, and old conversations will follow the new behaviour rather than a stale prompt.

The simplest form is a string:

instructions_static.py
from pydantic_ai import Agent

agent = Agent(
    'openai:gpt-5.2',
    instructions=(
        'You are a support assistant for an online bookshop. '
        'Answer politely and in at most three sentences. '
        'If you do not know, say so instead of guessing.'
    ),
)

Sometimes the instruction depends on information that only exists at run time, such as the name of the user asking. For that, decorate a function with @agent.instructions. It receives a RunContext, an object we describe in the dependencies section, and returns a string:

instructions_dynamic.py
from dataclasses import dataclass
from pydantic_ai import Agent, RunContext

@dataclass
class Deps:
    customer_name: str

agent = Agent('openai:gpt-5.2', deps_type=Deps)

@agent.instructions
def personalise(ctx: RunContext[Deps]) -> str:
    return f"The customer's name is {ctx.deps.customer_name}. Greet them by name."

result = agent.run_sync('Hi there', deps=Deps(customer_name='Mona'))
print(result.output)

You can combine a constructor string with one or more decorated functions.

Try it
  1. Give an agent the bookshop instructions above and ask it for the capital of France.
  2. Tighten the instructions so that it refuses off-topic questions politely, and ask again.
  3. Compare the two replies.
the same model behaving differently after only a sentence changed, which is the main lever you have.

Structured output: answers your code can trust

This is the feature that most beginners come to Pydantic AI for. Instead of a string you must parse, you declare the shape of the answer as a Pydantic model and the agent returns an instance of it.

A Pydantic model is a class whose fields have type annotations. Pydantic checks that data matches those types and converts it where it sensibly can. Pass it as output_type:

structured_output.py
from pydantic import BaseModel
from pydantic_ai import Agent

class City(BaseModel):
    city: str
    country: str

agent = Agent(
    'google:gemini-3-flash-preview',
    output_type=City,
    instructions='Answer with the city and country.',
)

result = agent.run_sync('Where were the 2012 Olympics held?')
print(result.output)
print(result.output.city)

The printed output is a City object, shown as City(city='London', country='United Kingdom'), and result.output.city is plain 'London'. Your editor knows result.output is a City because the agent is generic in its output type, so autocompletion and type checking work on the answer.

What happens if the model replies with something that does not fit, such as a missing field or a number where text was expected? Pydantic raises a validation error, the library turns that into a message to the model describing what was wrong, and the model gets another chance. You did not write any retry code. By default the output has a small retry budget; if the model still cannot comply, you get an UnexpectedModelBehavior error, and we discuss reading it in the errors section.

You can add richer rules with the usual Pydantic tools. Field descriptions help the model understand what you mean, and constraints reject nonsense:

structured_rules.py
from pydantic import BaseModel, Field
from pydantic_ai import Agent

class Review(BaseModel):
    sentiment: str = Field(description='one of: positive, negative, neutral')
    score: int = Field(ge=1, le=5, description='rating from 1 to 5')
    summary: str = Field(description='one sentence summary')

agent = Agent('openai:gpt-5.2', output_type=Review)
result = agent.run_sync('The delivery was late but the book was wonderful.')
print(result.output.score)

If the model returns a score of 9, validation fails on le=5 and the model is told to correct it. This is the heart of the library: the constraints you already know how to write in Pydantic become guardrails around the model.

Output types are not limited to classes. You can ask for a plain int, a list[str], a dataclass or a typed dictionary. When the answer could be one of several shapes, pass a list of types, which means "any one of these":

union_output.py
from pydantic import BaseModel
from pydantic_ai import Agent

class Box(BaseModel):
    width: int
    height: int

agent = Agent('openai:gpt-5.2', output_type=[Box, str])
result = agent.run_sync('A box that is 10 wide and 20 high')
print(result.output)

Here the model may return a Box, or a string if it cannot. Use the list form rather than Box | str, because the type checker mypy has trouble with the pipe form in this position.

A valid shape is not a true answer Validation guarantees the answer has the right form. It cannot guarantee the content is correct. A model can return a well-formed City that names the wrong city. Check anything important against a trusted source, and use output validators, introduced in the Mid-level guide, for rules that need your own code.
Try it
  1. Write a Review model like the one above and run it on three different sentences.
  2. Add a field with a constraint the model might break, such as max_length=20 on the summary.
  3. Run it on a long sentence and observe whether the agent recovers.
typed objects instead of strings, and, if you are lucky, a visible recovery where a constraint was broken and then satisfied.

Function tools: letting the model act

A model knows only what it was trained on and what you put in the prompt. It cannot know today's date, an order's status, or the contents of your database. A tool fixes that: you write a Python function, and the model can ask the agent to run it.

There are two decorators. Use @agent.tool_plain for a function that needs nothing from the outside, and @agent.tool for one that needs the run context, which gives access to your dependencies. Begin with the plain one:

tool_plain.py
import random
from pydantic_ai import Agent

agent = Agent(
    'openai:gpt-5.2',
    instructions="Use the tool to roll the die, then tell the player what came up.",
)

@agent.tool_plain
def roll_dice() -> str:
    """Roll a six-sided die and return the result."""
    return str(random.randint(1, 6))

result = agent.run_sync('Roll the die for me.')
print(result.output)

Walk through what happens. The library inspects roll_dice and tells the model there is a tool with this name and this description. The docstring becomes the description, so write it for the model: say what the function does and when to use it. The model replies with a request to call roll_dice; the library runs it and sends back the string it returned; the model then writes the final answer. The program prints a sentence like "You rolled a 4."

Tools can take arguments, and the type hints become the argument schema. Documenting the parameters in the docstring helps the model fill them in correctly. The library understands the common docstring styles (Google, NumPy and Sphinx):

tool_args.py
from pydantic_ai import Agent

agent = Agent('openai:gpt-5.2')

ORDERS = {'A100': 'shipped', 'A101': 'processing'}

@agent.tool_plain
def order_status(order_id: str) -> str:
    """Look up the status of a customer order.

    Args:
        order_id: the order identifier, such as A100
    """
    return ORDERS.get(order_id, 'unknown order')

result = agent.run_sync('Where is order A101?')
print(result.output)

The arguments the model supplies are validated against the signature before your function runs. If the model sends a number where a string is required, or omits a required argument, validation fails and the model is told what was wrong, so a bad call never reaches your code.

A tool can also be async def, and for anything that waits on a network or a database you should prefer that. Synchronous tools still work, since the library runs them in a thread pool, but async tools let many runs share one event loop efficiently.

Tools can fail on purpose. If a tool receives a value that is syntactically valid but wrong, such as an order that does not exist, raise ModelRetry with a message, and the model is asked to try again with that guidance:

tool_retry.py
from pydantic_ai import Agent, ModelRetry

agent = Agent('openai:gpt-5.2')

@agent.tool_plain
def order_status(order_id: str) -> str:
    """Look up the status of a customer order."""
    if not order_id.startswith('A'):
        raise ModelRetry('Order ids start with the letter A, for example A100.')
    return 'shipped'

The message you raise is sent back as if it were a tool result, so write it the way you would explain the mistake to a colleague. The default budget is a single retry per tool, which you can raise with @agent.tool_plain(retries=3) when a tool is genuinely hard to call correctly.

The docstring is part of the prompt When the model uses the wrong tool, or fills in an argument badly, the first fix is almost always a better docstring, not more code. Say what the tool is for, name the format of each argument with an example, and say when not to use it.
Give tools the least power you can A tool runs with whatever access your program has. A model that can call delete_record may well call it. Start with read-only tools, validate arguments yourself inside the tool, and for anything destructive use the approval feature described in the Mid-level guide rather than trusting the model's judgement.
Try it
  1. Add the order_status tool to an agent and ask about an order that exists and one that does not.
  2. Print result.all_messages() and find the tool call and the tool return in the list.
  3. Make the docstring worse, for example delete it, and observe whether the model still uses the tool well.
a visible trail of request, tool call, tool result and final answer, and a sense of how much the docstring matters.

Dependencies and RunContext

Tools usually need something: a database connection, an HTTP client, a user id. You could use global variables, but globals are hard to test and easy to get wrong when many requests run at once. Pydantic AI offers dependency injection instead, in the same spirit as FastAPI.

The recipe has three steps. Define a class that holds what your tools need, usually a dataclass. Tell the agent its type with deps_type. Then pass an instance in at run time with deps=, and read it inside tools through ctx.deps:

deps_example.py
from dataclasses import dataclass
from pydantic_ai import Agent, RunContext

class FakeDb:
    async def get_order(self, order_id: str) -> str:
        return {'A100': 'shipped', 'A101': 'processing'}.get(order_id, 'unknown')

@dataclass
class Deps:
    db: FakeDb
    customer: str

agent = Agent(
    'openai:gpt-5.2',
    deps_type=Deps,
    instructions='You are a helpful support assistant.',
)

@agent.tool
async def lookup_order(ctx: RunContext[Deps], order_id: str) -> str:
    """Look up an order's status.

    Args:
        order_id: the order identifier
    """
    status = await ctx.deps.db.get_order(order_id)
    return f'Order {order_id} for {ctx.deps.customer} is {status}.'

result = agent.run_sync(
    'Where is order A100?',
    deps=Deps(db=FakeDb(), customer='Mona'),
)
print(result.output)

Look at the first parameter of lookup_order: ctx: RunContext[Deps]. This is the run context, a small object the library passes to any function that wants it. The [Deps] part tells the type checker what ctx.deps is, so ctx.deps.db autocompletes and a typo is caught before you run anything. The context does not appear in the schema the model sees; the model only sees order_id.

An important detail: deps_type=Deps is only for type checking. It does not create the dependencies. You must still pass deps=Deps(...) on every run, and if a tool needs them and you forgot, you will get an AttributeError on ctx.deps at the first use rather than a helpful message.

Why go to this trouble? Three reasons. Tools stay free of hidden global state, so they are easy to read. Each request can carry its own user and its own connection, which is exactly what a web service handling many customers needs. And in tests you can pass a fake database, as the example does, and never touch the real one.

Try it
  1. Run the dependency example and change customer to your own name.
  2. Add a second tool that reads another field from Deps.
  3. Delete the deps= argument from the run call and read the error you get.
per-call data reaching your tools through the context, and a clear failure when you forget to supply it.

Conversations: carrying history between runs

A single run has no memory of earlier runs. Ask "Where were the 2012 Olympics held?" and then, in a separate call, "And its population?", and the second call has no idea what "its" refers to. To continue a conversation you pass the earlier messages back in.

Every result carries its messages. result.new_messages() returns only the messages from the run that just finished, and result.all_messages() returns everything including any history you passed in. The usual pattern is to hand the new messages of one run to the next:

conversation.py
from pydantic_ai import Agent

agent = Agent('openai:gpt-5.2', instructions='Be concise.')

first = agent.run_sync('Where were the 2012 Olympics held?')
print(first.output)

second = agent.run_sync(
    'And what is its population?',
    message_history=first.new_messages(),
)
print(second.output)

The second call receives the first exchange as context, so "its" resolves to London. Because instructions are re-evaluated each run and are not stored in history, the second run still obeys the current instructions.

To keep going for many turns, collect the messages as you go. A common beginner pattern is a simple loop:

chat_loop.py
from pydantic_ai import Agent

agent = Agent('openai:gpt-5.2', instructions='You are a friendly tutor.')
history = []

while True:
    question = input('You: ')
    if question.strip().lower() in {'quit', 'exit'}:
        break
    result = agent.run_sync(question, message_history=history)
    history = result.all_messages()
    print('Agent:', result.output)

Each turn, all_messages() already contains the earlier history plus the new turn, so assigning it back to history keeps the whole conversation.

There are two things to understand about history. First, it grows, and every message is sent to the model again on every turn, so long conversations cost more and eventually exceed the model's context window. Production systems trim or summarise old messages, a topic for the Mid-level guide. Second, history is yours to store. The agent keeps nothing between runs. If you want a conversation to survive a restart, save the messages, for example as JSON, and load them later. The library gives you a type adapter for that:

save_history.py
from pydantic_ai import ModelMessagesTypeAdapter
from pydantic_core import to_json

blob = to_json(result.all_messages())          # bytes you can store anywhere
history = ModelMessagesTypeAdapter.validate_json(blob)

Store the bytes in a file, a database column or a cache, and turn them back into messages with validate_json. Treat saved history as trusted data you produced yourself. If a browser or other untrusted client can send you history, the Senior guide explains why that is dangerous and how to clean it.

Try it
  1. Run the two-call conversation and confirm the second answer uses the first as context.
  2. Remove message_history from the second call and see how the answer changes.
  3. Save the messages with to_json to a file, restart Python, load them, and continue the chat.
a follow-up question that makes sense only with history, and a conversation that survives a restart.

Streaming: showing the answer as it arrives

Language models produce text a piece at a time, and waiting for the whole reply makes an application feel slow. Streaming lets you show text as it is generated. Streaming uses an async context manager, so it lives inside an async def:

streaming.py
import asyncio
from pydantic_ai import Agent

agent = Agent('openai:gpt-5.2', instructions='Be concise.')

async def main():
    async with agent.run_stream('Explain what an API is.') as response:
        async for chunk in response.stream_text():
            print(chunk, end='', flush=True)
    print()

asyncio.run(main())

By default stream_text() yields the text accumulated so far each time, so you may see the sentence grow rather than receive only the new piece. If you want only the new fragment each time, pass delta=True:

streaming_delta.py
async with agent.run_stream('Explain what an API is.') as response:
    async for piece in response.stream_text(delta=True):
        print(piece, end='', flush=True)

There is a trap in streaming that the documentation calls out. run_stream() treats the first output that matches your output type as the final answer. With the default settings, if the model writes some text before calling a tool, the stream can end at the text, and the tool call that follows never runs. If you need the full tool loop to continue while you watch events, use run_stream_events() or iter(), which the Mid-level guide covers. For a simple streaming chat with no tools, run_stream is exactly right.

Try it
  1. Run the streaming script and watch the answer appear.
  2. Switch between the default and delta=True and compare what is printed.
text appearing progressively, and a feel for the difference between accumulated text and increments.

Limits, retries and reading the errors

Agents loop, and loops need brakes. Pydantic AI ships sensible defaults, and understanding them turns scary tracebacks into short conversations.

The most important default is the request limit: a single run may make at most 50 model requests. A run that keeps calling tools in circles will stop with a UsageLimitExceeded error whose message begins "The next request would exceed the request_limit of 50". You can set your own limits with UsageLimits:

limits.py
from pydantic_ai import Agent, UsageLimits, UsageLimitExceeded

agent = Agent('openai:gpt-5.2')

try:
    result = agent.run_sync(
        'Write a long essay about rivers.',
        usage_limits=UsageLimits(output_tokens_limit=200),
    )
except UsageLimitExceeded as exc:
    print('Stopped:', exc)

You can limit requests, tool calls, input tokens, output tokens, total tokens and an approximate dollar cost. Treat the cost limit as a guide, not a billing guarantee, since it depends on bundled price data. Pair it with a spend cap set at your provider.

The next default is the retry budget, which is one. A tool or an output that fails validation once gets one more chance; if it fails again the run stops. Raise it on the agent with Agent(..., retries=3), or per tool with @agent.tool(retries=3).

Now the errors, one by one, because you will meet all of them.

Message What it means What to do
UserError: Set the OPENAI_API_KEY environment variable... The provider cannot find a key Export the variable named in the message
UserError: Unknown model: gpt-5 No provider prefix, or an unknown prefix Write openai:gpt-5
ImportError: Please install the ... package The extra for that provider is not installed pip install "pydantic-ai[groq]" or similar
UnexpectedModelBehavior: Tool 'x' exceeded max retries count of 1 The model kept calling a tool badly Improve the docstring, or raise retries
UnexpectedModelBehavior: Exceeded maximum output retries The answer never satisfied your output type Simplify the type, relax constraints, or raise output retries
UsageLimitExceeded: The next request would exceed the request_limit of 50 A loop hit the default cap Find the loop, or set a deliberate limit
ModelHTTPError The provider returned an error such as 401, 404 or 429 Check the key and model name; for 429, slow down
RuntimeError: This event loop is already running run_sync inside a notebook Use await agent.run(...)

When a run fails in a way you cannot explain, the single most useful tool is capture_run_messages. It records every message exchanged, even when the run raises:

debug_messages.py
from pydantic_ai import Agent, capture_run_messages, UnexpectedModelBehavior

agent = Agent('openai:gpt-5.2', output_type=int)

with capture_run_messages() as messages:
    try:
        agent.run_sync('Name a colour.')
    except UnexpectedModelBehavior:
        for message in messages:
            print(message)

Reading those messages shows you what the model actually said and what validation told it. Nine times in ten, the cause is visible there: the model returned text when you asked for an integer, or called a tool with the wrong argument name.

An error that surprises people after the move to version 2 is a TLS certificate failure, SSL: CERTIFICATE_VERIFY_FAILED. The library's own web requests now verify certificates against the operating system's trust store instead of a bundled list. In a slim container image, or behind a corporate proxy that inspects encrypted traffic and re-signs it with its own certificate authority, that authority may be missing from the system store. The fix is to install the certificate into the operating system's store, or to pass your own configured HTTP client to the provider.

An error is information, not a crash to hide Do not wrap runs in a bare except Exception: pass. Catch the specific exceptions you can handle, such as UsageLimitExceeded or ModelHTTPError, and let the rest surface. The message text almost always names the fix.
Try it
  1. Set UsageLimits(output_tokens_limit=10) and run a prompt that needs a longer answer.
  2. Read the UsageLimitExceeded message and find the limit and the observed count in it.
  3. Use capture_run_messages around a run that fails and print the messages.
a limit that fires on purpose, and a printed transcript of what the model saw and said.

Testing without a model

An agent that calls a paid, unpredictable model is awkward to test. Pydantic AI solves this with test models that run locally, so your test suite is fast, free and repeatable.

TestModel is a stand-in that calls your tools with generated arguments and returns a valid placeholder answer. You swap it in with agent.override, which affects only the code inside the with block:

test_agent.py
from pydantic_ai import Agent, models
from pydantic_ai.models.test import TestModel

models.ALLOW_MODEL_REQUESTS = False   # fail loudly if anything reaches a real model

agent = Agent('openai:gpt-5.2')

@agent.tool_plain
def order_status(order_id: str) -> str:
    """Look up an order."""
    return 'shipped'

def test_agent_runs():
    with agent.override(model=TestModel()):
        result = agent.run_sync('Where is my order?')
    assert result.output is not None

The line models.ALLOW_MODEL_REQUESTS = False is a safety net. If any test accidentally reaches a real model, it raises RuntimeError: Model requests are not allowed, since ALLOW_MODEL_REQUESTS is False instead of quietly spending money. When you want a model whose replies you script yourself, FunctionModel lets you pass a function that returns whatever response you like.

Try it
  1. Write the test above and run it with pytest.
  2. Remove the override block while keeping ALLOW_MODEL_REQUESTS = False and read the error.
a passing test that costs nothing, and a loud failure the moment a test would call a real model.

The clai command line

Installing Pydantic AI also gives you clai, a command-line chat client. It is the quickest way to talk to a model from a terminal and a good way to check that your key and model name work before you write code.

BASH
clai

With no arguments it starts an interactive session. Type a question and press Enter. A prompt in quotes runs one question and exits:

BASH
clai "Explain DNS in two sentences"

Choose a model with -m, and list the model names the library knows with -l:

BASH
clai -m anthropic:claude-sonnet-4-6 "Explain DNS in two sentences"
clai --list-models

If you do not want streamed output, add --no-stream. Inside an interactive session, the slash commands include /exit to leave, /markdown to switch markdown rendering, /multiline for multi-line input, /cp to copy the last answer, and /usage to see token usage.

You can also run clai against your own agent. The -a option takes module:variable, so if support.py defines agent = Agent(...), then this opens a chat with it:

BASH
clai -a support:agent

That is useful for trying instructions and tools by hand without writing a chat loop. clai also has a web command that opens a browser chat, which is convenient, but it is for local development only. Never expose it to a network. Earlier versions of the web interface had security weaknesses that were patched in releases 2.28 and 2.30, and even a patched development server is not designed to face users.

Try it
  1. Run clai, ask a question, then use /usage.
  2. Run clai -a against your bookshop agent from earlier.
a working chat in your terminal, and a quick way to poke at an agent without any extra code.

Configuration and everyday settings

Beyond instructions and tools, a few settings shape how a model behaves, and you pass them as model_settings. They are a plain dictionary:

settings.py
from pydantic_ai import Agent

agent = Agent(
    'openai:gpt-5.2',
    model_settings={'temperature': 0.0, 'max_tokens': 500},
)

result = agent.run_sync(
    'Summarise: Pydantic AI is a Python agent library.',
    model_settings={'max_tokens': 100},
)

The common fields are temperature, which controls randomness (a low value makes answers more repeatable), max_tokens, which caps the length of the reply, top_p, timeout, seed and parallel_tool_calls. Settings can be given when you build the model, on the agent, and on each run, and later layers win, so a run-level value overrides the agent's. Not every provider supports every field, so if a setting seems ignored, check that provider's page in the documentation.

Two other agent arguments are worth knowing at this level. name labels the agent, which helps in logs. retries sets the default retry budget, as discussed. There are many more, such as toolsets, capabilities and end_strategy, but you can leave them alone until you have a reason.

For observability, which means seeing what your agents do, Pydantic AI emits standard OpenTelemetry traces and works with Logfire, the company's own service, as well as with any OpenTelemetry-compatible backend. The simplest setup with Logfire is two lines after you authenticate:

tracing.py
import logfire

logfire.configure()
logfire.instrument_pydantic_ai()

After that, each run shows up as a trace you can inspect, including every model call and tool call. If you already use a tracing tool for language models, the guide on Langfuse shows how traces help you understand and debug agent behaviour. Tracing is optional for a beginner, but turning it on early makes mysterious behaviour much easier to explain.

Prompts in traces can contain private data Traces record what was sent to the model, which may include customer details. If your data is sensitive, check what your tracing backend stores and where it lives, and use the library's setting that excludes message content from traces. This matters for any employer with data-protection obligations.
Try it
  1. Run the same prompt five times with temperature at 0.0 and five times at a high value.
  2. Compare how much the answers vary.
repeatable answers at low temperature and more variety at high temperature, which is the practical meaning of the setting.

Recognising version 1 code

Because Pydantic AI moved to version 2 in June 2026, tutorials, forum answers and AI-generated snippets often show version 1 code. Recognising it saves hours. Here are the differences a beginner is most likely to hit.

You see this (version 1) Write this (version 2)
Agent('gpt-5') with no prefix Agent('openai:gpt-5')
result.usage() result.usage
result.data (very old) result.output
Agent(..., mcp_servers=[...]) Agent(..., toolsets=[...])
MCPServerStdio, MCPServerSSE and friends pydantic_ai.mcp.MCPToolset
Agent(..., builtin_tools=[...]) Agent(..., capabilities=[NativeTool(...)])
google-gla: or google-vertex: prefixes google: or google-cloud:
result.stream(...) on a streamed result stream_output or stream_text

The theme of version 2 is that many options once passed to the Agent constructor moved onto capabilities, reusable bundles of tools and behaviour that you pass as capabilities=[...]. You will not need them yet, but you will see the word often in current documentation. If you do inherit a version 1 codebase, the official advice is to upgrade to the latest 1.x release first, fix every deprecation warning, and only then move to version 2. Messages you saved with version 1 still load in version 2.

Try it
  1. Find a Pydantic AI snippet on the web that is more than six months old.
  2. List every line in it that matches the left column above.
  3. Rewrite it for version 2 and run it with a test model.
practice at spotting stale patterns, which is a real skill when a library changes this fast.

Pin the version in anything you deploy, for example pydantic-ai==2.52.*, because the project releases every day or two.

Putting it all together: a small support assistant

Let us combine everything into one program: typed output, a tool that reads dependencies, instructions, history, limits and a test. The scenario is a bookshop assistant that answers questions about orders and always returns a structured reply.

assistant.py
import asyncio
from dataclasses import dataclass

from pydantic import BaseModel, Field
from pydantic_ai import Agent, ModelRetry, RunContext, UsageLimits


class OrderDb:
    """A stand-in for a real database client."""

    ORDERS = {
        'A100': {'status': 'shipped', 'eta_days': 2},
        'A101': {'status': 'processing', 'eta_days': 5},
    }

    async def find(self, order_id: str) -> dict | None:
        return self.ORDERS.get(order_id)


@dataclass
class Deps:
    db: OrderDb
    customer: str


class Reply(BaseModel):
    answer: str = Field(description='a polite answer of at most three sentences')
    needs_human: bool = Field(description='true if a person should follow up')


agent = Agent(
    'openai:gpt-5.2',
    deps_type=Deps,
    output_type=Reply,
    retries=2,
)


@agent.instructions
def brief(ctx: RunContext[Deps]) -> str:
    return (
        'You are a support assistant for an online bookshop. '
        f'You are speaking to {ctx.deps.customer}. '
        'Use the lookup tool for any question about an order. '
        'If the order cannot be found, set needs_human to true.'
    )


@agent.tool
async def lookup_order(ctx: RunContext[Deps], order_id: str) -> str:
    """Look up an order by its identifier.

    Args:
        order_id: the order identifier, for example A100
    """
    if not order_id.startswith('A'):
        raise ModelRetry('Order ids start with the letter A, for example A100.')
    order = await ctx.deps.db.find(order_id)
    if order is None:
        return 'No such order.'
    return f"Status {order['status']}, arriving in {order['eta_days']} days."


async def main() -> None:
    deps = Deps(db=OrderDb(), customer='Mona')
    limits = UsageLimits(request_limit=10, total_tokens_limit=5000)
    history = []

    for question in ['Where is order A100?', 'And when will it arrive?']:
        result = await agent.run(
            question,
            deps=deps,
            message_history=history,
            usage_limits=limits,
        )
        history = result.all_messages()
        print(result.output)


if __name__ == '__main__':
    asyncio.run(main())

Read it from the top. OrderDb is the thing your tools need, and Deps bundles it with the customer's name. Reply is the shape of every answer, with a boolean so that calling code can route hard cases to a person. The agent is created once, with a retry budget of two.

The brief function builds the instructions at run time, using the customer's name from the dependencies. The lookup_order tool does the real work: it rejects a malformed id with ModelRetry, which sends the guidance back to the model, reads the database through ctx.deps, and returns a short string.

In main, the same deps object and the same limits are reused for both questions, and history carries the first exchange into the second, which is why "it" in "when will it arrive?" makes sense. The limits make sure that one runaway run cannot spend more than ten requests or five thousand tokens.

Now a test, so a change to the code is checked without a model:

test_assistant.py
from pydantic_ai import models
from pydantic_ai.models.test import TestModel

from assistant import Deps, OrderDb, agent

models.ALLOW_MODEL_REQUESTS = False


def test_assistant_returns_a_reply():
    with agent.override(model=TestModel()):
        result = agent.run_sync(
            'Where is order A100?',
            deps=Deps(db=OrderDb(), customer='Mona'),
        )
    assert isinstance(result.output.needs_human, bool)

Run it with pytest. It proves the wiring: the tool registers, the dependencies are read, the output type is satisfied. Then run python assistant.py with a real key to see real behaviour.

Try it
  1. Run the assistant with a real key and ask about an order that does not exist.
  2. Check that needs_human comes back true.
  3. Add a third question to the loop that depends on the second answer.
a multi-turn, tool-using, typed assistant that you built from the five nouns, with a test around it.

What you can now do, and what comes next

You can now install Pydantic AI, check that it works with no key, and talk to a real model through a provider-prefixed name. You can write instructions, ask for typed output as a Pydantic model, and rely on validation to keep the model honest. You can give the model tools, pass dependencies through the run context, hold a conversation by passing messages, stream a reply, set limits, read the common errors, and test the lot with TestModel. You can also spot version 1 code and translate it.

A short recap of the questions you should be able to answer without looking:

Can you... The one-line answer
Name the five nouns? Agent, model, tool, output, dependencies
Pick a model? A provider:model string such as openai:gpt-5.2
Get a typed answer? output_type= a Pydantic model
Give the model a function? @agent.tool_plain or @agent.tool with a good docstring
Pass a database to a tool? deps_type, then deps=, then ctx.deps
Continue a conversation? message_history=previous.new_messages()
Read a run's token usage? result.usage, a property
Test with no model? agent.override(model=TestModel())
Stop a runaway loop? UsageLimits, and the default 50 request cap

Mid-level takes every one of those topics further: the choice between output modes and when to use each, output validators, how retries really layer, human approval for risky tools, managing and trimming history, the five run methods in depth, connecting external tools through MCP, evaluating agent quality, and capabilities and hooks. If your agents will sit behind a web service, the guide on FastAPI covers the other half of that stack, and the guide on the Model Context Protocol explains the tool standard Pydantic AI can plug into.

Senior then covers what you own when agents are a platform for a team: the trust model for client-supplied history and approvals, secrets and per-user credentials, multi-tenant isolation, concurrency and rate limiting, fallback across providers, durable execution so a run survives a crash, cost control, version upgrades, and the cases where a different tool is the better choice.

Sources