تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
OpenAI Agents SDKLLMsFrameworks & agents3 مستويات95 قسمًايغطّي openai-agents 0.22دليل بالإنجليزية

The Complete OpenAI Agents SDK Guide

Build agents with tools, handoffs, guardrails and tracing on OpenAI’s lightweight agent SDK. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
17sections
24examples

This is part one of three. It covers everything you need to build a working agent with the OpenAI Agents SDK, not a teaser. By the end you can define an agent, give it Python functions it can call, make it return a typed object instead of loose prose, route a question to a specialist agent, keep a conversation going across turns, stream tokens to a user, stop a request you do not want to serve, and read the stack trace when any of that goes wrong. The Mid-level and Senior guides take the same topics further into production; nothing you learn here is thrown away.

Every section ends with a Try it task. Do them as you go. Agents are one of those subjects where the vocabulary sounds obvious and the behaviour is not, and the gap closes only once you have watched your own agent call a tool, hand off to another agent, and loop ten times before giving up.

The version this guide was verified against is openai-agents 0.22.3, released 17 September 2026. The SDK is young and moves fast, and some of what you will read in older blog posts is now wrong, so a few places in this guide say explicitly "this changed".

What the OpenAI Agents SDK is, and the problem it solves

An agent in this context is a language model that has been given a job description, a set of actions it can take, and permission to decide for itself which action to take next. That last part is what separates an agent from a single API call. When you call a chat completion endpoint directly, you get one response and you are done. When you run an agent, the model may answer immediately, or it may decide to look something up first, call a function you wrote, read the result, call another function, and only then answer.

Somebody has to run that back-and-forth. Something has to take the model's request to call get_order_status, find the matching Python function, validate the arguments the model invented, call the function, catch the exception if it throws, put the result back into the conversation in the exact shape the API expects, call the model again, and decide when the whole thing is finished. That loop is not conceptually hard, but it is a few hundred lines of fiddly code that every team was writing from scratch and getting slightly wrong.

The OpenAI Agents SDK is that loop, written once, as a small Python library. You describe the agent; the SDK runs the loop. It is the production successor to Swarm, an earlier experimental OpenAI project, and by default it talks to the OpenAI Responses API and wraps it in a runtime.

AGENTinstructions + tools
→
RUNNERruns the loop
→
TOOLSyour Python functions
→
RESULTfinal_output

The SDK's own documentation gives a clean rule for when to use it at all, and it is worth internalising before you write a line of code. Call the Responses API directly if you want to own the loop, the tool dispatch and the state yourself — if you are building something unusual and the loop is the interesting part of your product. Use the SDK if you want the runtime to handle turns, tools, guardrails, handoffs, sessions or sandboxes, so that the interesting part of your product can be your tools and your prompts.

Two design decisions shape everything that follows, and both are easier to accept now than to discover by accident later.

It is a library, not a server. There is no agent daemon to install, no control plane, no dashboard you deploy. The SDK runs inside your own Python process — a FastAPI handler, a Celery worker, a script on your laptop. This is why there is nothing to configure beyond environment variables, and why questions like "where does conversation history live?" have the answer "wherever you decide to put it".

The primitive set is deliberately tiny. There are only a handful of concepts: agents, tools, handoffs, guardrails, sessions, and a runner that ties them together. Frameworks in this space often ship dozens of abstractions; this one ships about six, on the theory that you can read all of them in an afternoon and then predict what your program will do. That bet mostly pays off, and the small vocabulary is the main reason the SDK is a good first agent framework even if you later move to something else.

Try it
  1. Write down one task at your work that a language model could do if it could call two or three functions you already have.
  2. Name those functions and write their Python signatures, with type hints, on paper.
  3. Write one sentence describing what the model should do, as if you were briefing a new colleague.
you now have, almost literally, the inputs to an Agent(...) call: the sentence becomes instructions and the signatures become tools. Keep the paper; you will build this agent by the end of the guide.

The four nouns: Agent, Runner, Tool, Handoff

Nearly every sentence in the SDK documentation uses one of four words in a precise sense. Getting them straight now saves a great deal of confusion.

An Agent is a configured LLM. You create one with Agent(name=..., instructions=..., tools=..., ...). The instructions field is what other frameworks call the system prompt: the standing brief the model sees on every call. name is not cosmetic — it appears in traces, and when one agent hands off to another the handoff tool is named after it. An agent also optionally carries handoffs, guardrails, an output_type, a model name and model settings. An agent object is cheap and inert; creating one makes no network call.

A Runner is what actually executes an agent. It is a class with three entry points you will use constantly: Runner.run(...) which is async and is what you should reach for by default, Runner.run_sync(...) which is a synchronous wrapper for scripts and notebooks-that-are-not-notebooks, and Runner.run_streamed(...) which returns a streaming result you iterate over. All three run the same loop.

A Tool is anything the model can call. The kind you will write first is a function tool: an ordinary Python function with a decorator on it. The SDK reads the function's name, its docstring and its type hints, and from those builds the JSON schema the model needs in order to call it correctly. There are other kinds — tools hosted by OpenAI such as web search, local runtime tools such as a shell, tools exposed by an MCP server, and even other agents wrapped as tools — but a decorated function is the one that will do ninety percent of your work.

A Handoff is delegation in which the other agent takes over the conversation. The model sees a handoff as a tool named transfer_to_<agent_name>, and when it calls that tool the runner swaps in the target agent and keeps looping. Handoffs stay inside a single run, which means one call to Runner.run can begin with a triage agent and finish with a billing specialist.

The distinction you must not blur is handoff versus agent-as-tool, because both look like "one agent uses another agent" and they behave very differently. With a handoff, control and the conversation move to the specialist, and when the run finishes result.last_agent is the specialist, not the agent you started with. With an agent-as-tool — created by calling agent.as_tool(tool_name=..., tool_description=...) — the orchestrating agent calls the specialist the way it would call any function, gets a string back, and keeps control. The nested run receives input the orchestrator generated for it, and its state does not inherit the parent's conversation history unless you deliberately pass the same session.

Two more nouns you will meet shortly A session is a memory layer that loads conversation history before each run and saves the new items after it. A guardrail is a validation function that can stop a run before or after the agent does its work. Both are optional, and both are covered later in this guide.
Try it
  1. Take your task from the previous section and decide whether it needs one agent or two.
  2. If two, decide whether the second should be a handoff or an agent-as-tool, and write one sentence justifying it.
  3. Ask yourself who should be talking to the user at the end of the run. That answer decides it.
"who finishes the conversation" is the whole test. If the specialist should finish it, use a handoff. If your original agent should finish it, use as_tool.

The agent loop, turn by turn

Everything the runner does is one loop with three exits. It is short enough to memorise, and memorising it makes the SDK's behaviour predictable instead of magical.

  1. Call the modelThe runner calls the LLM for the current agent, with the current input: the instructions, the conversation so far, and the schemas of every tool the agent has.
  2. Is the output final?If the model returned text of the desired output_type and made no tool calls, the run is over and that text is the result. Both halves matter: text accompanied by a tool call is not final.
  3. Is it a handoff?If the model called a transfer_to_* tool, the runner switches the current agent to the target and loops back to step one.
  4. Are there tool calls?If so, the runner executes them, appends their results to the conversation, and loops back to step one.
  5. Have we looped too long?If the number of model calls exceeds max_turns, the runner raises MaxTurnsExceeded.

The word that trips up almost every newcomer is turn. A turn is one model call, not one user message. If a user asks a single question and the agent calls three tools before answering, that is four turns, not one. max_turns defaults to 10, so an agent that gets stuck calling the same tool repeatedly will blow up after ten model calls rather than running forever and charging you for it. You can raise the number, and since version 0.16.0 you can pass max_turns=None to remove the limit entirely — which you should do only when you have some other way of stopping the loop.

The result object you get back has a few fields you will reach for immediately. result.final_output holds the answer: a str if the last agent had no output_type, otherwise an instance of that type. It is typed as Any, which looks sloppy until you remember that a handoff can change which agent finishes the run, so the static type genuinely is not knowable. result.last_agent tells you which agent produced the answer. result.to_input_list() gives you the whole conversation as a list of input items, which is how you do multi-turn chat by hand. And result.context_wrapper.usage carries token counts, which is how you find out what a run cost.

One subtlety about final_output: it can be None. That happens when the run paused rather than finished — most commonly because a tool needs human approval. That is a Senior-level topic, but it is worth knowing that None means "not finished", not "the model said nothing".

Reasoning tokens are turns too Modern reasoning models can spend a great deal of thinking before they emit a tool call. Each of those model calls is a turn, and each one costs tokens. If you see a run take far longer or cost far more than you expected, the first thing to look at is how many turns it actually used, and the second is the reasoning effort you asked for.
Try it
  1. On paper, trace the loop for this request: "What's the weather in Riyadh and should I take a coat?" against an agent with a single get_weather tool.
  2. Count the turns.
  3. Now trace it for an agent whose instructions say "always confirm the city with the user first".
the first is two turns: one to call the tool, one to answer. The second is three or more, because the confirmation round-trip adds a model call. Instructions change turn counts, and turn counts are the bill.

Installing the SDK and checking your setup

The package is named openai-agents and you import it as agents. Getting those two names mixed up produces a confusing ModuleNotFoundError, so note it once: install openai-agents, import agents. Python 3.10 or newer is required; support for 3.9 was dropped in version 0.9.0, and the current release is tested on 3.10 through 3.14.

On macOS or Linux:

BASH
mkdir my_project && cd my_project
python -m venv .venv
source .venv/bin/activate
pip install openai-agents          # or: uv init && uv add openai-agents
export OPENAI_API_KEY=sk-...

On Windows the activation step differs and the environment variable is set differently depending on your shell:

POWERSHELL
python -m venv .venv
.venv\Scripts\activate
pip install openai-agents
$env:OPENAI_API_KEY = "sk-..."
CMD
set "OPENAI_API_KEY=sk-..."

The virtual environment is not optional busywork. The SDK pins its dependencies fairly tightly — at 0.22.3 it requires openai 3.x and pydantic 2.x — and installing it into a system Python alongside an older openai library is the single most common way to end up with an import error that has nothing to do with your code.

Check that the install worked before you write anything:

BASH
python -c "import agents, importlib.metadata as m; print(m.version('openai-agents'))"
pip show openai-agents

The first command should print 0.22.3. If it prints an older version, you have an old pin somewhere; if it raises ModuleNotFoundError: No module named 'agents', your shell is using a different Python than the one you installed into, which is almost always a forgotten source .venv/bin/activate.

Then make a real call, which is the only way to confirm that your API key works:

smoke_test.py
from agents import Agent, Runner

agent = Agent(name="Assistant", instructions="You are a helpful assistant")
result = Runner.run_sync(agent, "Write a haiku about recursion in programming.")
print(result.final_output)

Three lines, one network call, and a haiku. If that prints, you are set up. If it raises an error about a missing API key, the key is not in the environment of the process you just ran — check echo $OPENAI_API_KEY in the same terminal.

The SDK ships a number of optional extras for things you do not need on day one, and it is useful to know they exist so you recognise them in other people's code: voice, realtime, viz for drawing agent graphs with Graphviz, redis, sqlalchemy, mongodb, docker, litellm and any-llm for non-OpenAI models, and a set of hosted sandbox providers. Install them with bracket syntax, and always quote the brackets, because zsh on macOS will otherwise try to expand them as a glob pattern and fail with a cryptic no matches found:

BASH
pip install "openai-agents[voice]"
pip install "openai-agents[redis,sqlalchemy]"
pip install "openai-agents[viz]"        # also needs the Graphviz system binary
Pin the minor version The SDK uses 0.Y.Z versions, and the minor number goes up for breaking changes to any non-beta public API. The patch number covers fixes, new features and changes to beta features. The official advice is to pin as openai-agents~=0.22.3 or >=0.22,<0.23 so that a routine pip install -U cannot break your agent.
Try it
  1. Create the virtual environment, install the SDK and run both verification commands.
  2. Run the haiku script.
  3. Deliberately break it: open a new terminal without the key set and run the script again. Read the error.
reading the missing-key error once, on purpose, in a calm moment, means you will recognise it instantly at 2am in a deployment that forgot to mount a secret.

Your first agent, properly

The haiku script used Runner.run_sync because it is the shortest thing that works. Real code should use the async form, so write that version next and get used to it:

first_agent.py
import asyncio
from agents import Agent, Runner

agent = Agent(
    name="History Tutor",
    instructions="You answer history questions clearly and concisely.",
)

async def main():
    result = await Runner.run(agent, "When did the Roman Empire fall?")
    print(result.final_output)

asyncio.run(main())

There is a trap in run_sync that is worth understanding properly, because it accounts for a large share of the confused questions newcomers ask. Runner.run_sync works by starting an event loop, running the async code inside it, and returning the result. If an event loop is already running in your process, it cannot do that, and it raises:

TEXT
AgentRunner.run_sync() cannot be called when an event loop is already running.

The places you will meet this are Jupyter notebooks, FastAPI route handlers, discord.py or aiogram bot callbacks, and anywhere else inside async def. The fix is always the same: use await Runner.run(...) instead. run_sync is for scripts and tests, not for servers.

Try it
  1. Write first_agent.py and run it.
  2. Now paste the same code into a Jupyter cell, using Runner.run_sync, and watch it fail.
  3. Fix the cell by using await Runner.run(...) directly, with no asyncio.run.
inside a notebook the loop is already running, so await works at the top level of a cell and run_sync does not. This is the single most reported beginner error with the SDK.

Giving the agent tools

An agent with no tools can only talk. Tools are how it does things, and in this SDK a tool is an ordinary Python function with a decorator:

weather_agent.py
import asyncio
from agents import Agent, Runner
from agents.decorators import tool

@tool
def get_weather(city: str) -> str:
    """Return weather info for the specified city."""
    return f"The weather in {city} is sunny"

agent = Agent(
    name="Weather",
    instructions="Use tools when helpful.",
    tools=[get_weather],
)

async def main():
    result = await Runner.run(agent, "What should I wear in Cairo today?")
    print(result.final_output)

asyncio.run(main())

Read what the decorator did, because it did more than it looks. The tool name comes from the function name, so the model sees a tool called get_weather. The description comes from the docstring, parsed with the griffe library, which is why that one-line docstring is not decoration — it is the only explanation the model gets of what the tool is for. The parameter schema comes from the type hints, converted to JSON Schema by Pydantic, which is why city: str matters and a bare city would not work nearly as well.

The practical consequence is that the quality of your tool docstrings and type hints is the quality of your agent's tool use. If the model keeps calling a tool with nonsense arguments, your first move should not be prompt surgery on the instructions; it should be rewriting the docstring to say what the parameter means, and tightening the type hint so the schema constrains it.

You will see two spellings of the decorator in the wild. @function_tool is the original, and @tool was added in version 0.19.0 as an alias in the new agents.decorators module. Neither is deprecated, both do exactly the same thing, and the current documentation uses @tool. This guide uses @tool; if you are reading older tutorials and see from agents import function_tool, it is the same decorator.

A handful of options on the decorator are worth knowing early, even if you do not use them today:

Option Default What it does
name_override None Shows the model a different tool name than the function name
description_override None Replaces the docstring-derived description
strict_mode True Keeps the JSON schema strict so the model cannot invent fields
failure_error_function None Controls what the model sees when your tool raises
is_enabled True A bool or (ctx, agent) callable that hides the tool from the model
timeout None Seconds, async tools only, with timeout_behavior to choose the outcome

The failure_error_function behaviour is counter-intuitive and worth pinning down. If you leave it alone, an exception inside your tool does not crash the run: the SDK catches it and sends the model the text "An error occurred while running the tool. Please try again.", and the loop continues. If you pass failure_error_function=None explicitly, the exception is re-raised and your run fails loudly. Leaving it out and passing None are therefore different things, which is a genuinely surprising API, and the distinction matters the first time you want a failing database call to fail your request instead of being quietly explained to the model.

Both synchronous and asynchronous functions work as tools. Since version 0.8.0 synchronous tools are run in worker threads, so a slow blocking call will not freeze the event loop — but it will still occupy a thread, so genuinely I/O-heavy tools are better written async def.

Finally, a tool can reach your application's own data. If the first parameter is annotated ctx: RunContextWrapper[T], the SDK passes in the context object you handed to Runner.run(..., context=obj), and your tool reads it at ctx.context. The crucial property of that context is that it is never sent to the LLM. It is the right place for a database handle, the logged-in user's ID, or a tenant identifier — things your tools need and the model has no business seeing.

Try it
  1. Write two tools: get_order_status(order_id: str) -> str returning hard-coded data, and convert_currency(amount: float, from_ccy: str, to_ccy: str) -> float with a fixed rate.
  2. Give both to one agent and ask it a question that needs both.
  3. Now delete one docstring and re-run the same question.
with no docstring the model becomes noticeably worse at choosing and populating that tool, because you removed the only description it had. This is the fastest demonstration that docstrings are load-bearing code here.

Structured output you can trust

Most of the time you do not actually want prose back. You want a record you can put in a database, and parsing it out of a paragraph with a regular expression is how you end up with a 3am incident. The SDK handles this with output_type: give the agent a Pydantic model, and result.final_output becomes an instance of that model.

extract_event.py
import asyncio
from pydantic import BaseModel
from agents import Agent, Runner

class CalendarEvent(BaseModel):
    name: str
    date: str
    participants: list[str]

agent = Agent(
    name="Calendar extractor",
    instructions="Extract calendar events",
    output_type=CalendarEvent,
)

async def main():
    result = await Runner.run(
        agent,
        "Aya and Omar are meeting about the Q4 roadmap on 12 November.",
    )
    event = result.final_output
    print(event.name, event.date, event.participants)

asyncio.run(main())

Under the hood the SDK converts your model to a JSON schema and asks the API to constrain the model's output to it. The validation happens before you see the result, so if final_output is a CalendarEvent, the fields really are there and really are the declared types. This is a much stronger guarantee than "I asked the model to reply in JSON and it usually does".

Remember from the loop that the definition of a final output is "text of the desired output_type and no tool calls". With an output_type set, an agent can still call tools as many times as it likes; the structured object is only produced on the turn where it decides it is finished.

Strict schemas are the default, and they come with real constraints that produce an error message you should learn to recognise. A strict JSON schema must have a non-nullable object at its root. So a free-form dict[str, Any] as your output_type, or a Union of two models at the top level, will be rejected with a UserError that says something like "The root of a strict JSON schema must be a non-nullable object" or "must not use anyOf". The fix is nearly always to wrap the thing you wanted in a Pydantic model with named fields — which is better design anyway, because now the model knows what you are asking for.

Works, and models well

  • A BaseModel with named, typed fields
  • Nested BaseModels for structure
  • list[Item] wrapped in a model: class Result(BaseModel): items: list[Item]
  • Optional fields declared explicitly

Rejected or unreliable

  • dict[str, Any] at the root
  • A bare Union or list as output_type
  • Fields explained only in the instructions, not the model
  • Free prose that you then parse yourself
Try it
  1. Add a location: str | None field to CalendarEvent and re-run with a sentence that has no location.
  2. Change output_type to dict and read the error you get.
  3. Change it back, and add a nested class Attendee(BaseModel) with a name and an email.
step two is the lesson: strict mode refuses a free-form dict, and the error tells you exactly why. Step three shows that nesting is fine, so you rarely need the dict.

Handing off to a specialist

One agent with twenty tools and a two-page brief behaves worse than three agents with clear jobs. Handoffs are how you split the work.

triage.py
import asyncio
from agents import Agent, Runner, handoff
from agents.extensions.handoff_prompt import RECOMMENDED_PROMPT_PREFIX

history_tutor = Agent(
    name="History Tutor",
    handoff_description="Specialist agent for historical questions",
    instructions="You answer history questions clearly and concisely.",
)

math_tutor = Agent(
    name="Math Tutor",
    handoff_description="Specialist agent for math questions",
    instructions="You explain math step by step.",
)

triage = Agent(
    name="Triage Agent",
    instructions=(
        f"{RECOMMENDED_PROMPT_PREFIX}\n"
        "Route each homework question to the right specialist."
    ),
    handoffs=[history_tutor, handoff(math_tutor)],
)

async def main():
    result = await Runner.run(triage, "Who was the first president?")
    print(result.last_agent.name)
    print(result.final_output)

asyncio.run(main())

Three details in that snippet are doing specific work.

handoff_description is the text the triage agent sees when deciding where to route. It is to a handoff what a docstring is to a tool: the only description available. An agent that routes badly usually has vague handoff descriptions, not a bad router prompt.

RECOMMENDED_PROMPT_PREFIX is a block of boilerplate the SDK ships that explains the handoff mechanism to the model — essentially, "you are part of a multi-agent system, here is how transferring works". Models route noticeably better with it, and prepending it costs you one import.

The handoffs list accepts either an agent directly or handoff(agent). They are equivalent until you need options, and handoff() is where the options live: tool_name_override if you dislike the generated transfer_to_math_tutor, tool_description_override, an on_handoff callback that fires when the transfer happens, an input_filter that lets you trim the history the specialist receives, and is_enabled to turn a route off dynamically. Start with the bare agent and reach for handoff() when you need one of those.

Because result.final_output is produced by whichever agent finished, an output_type set only on your triage agent will not apply when a specialist answers. If every route must return the same structured shape, set output_type on every agent that can finish a run. This catches people out regularly.

Try it
  1. Run the triage script with a history question, a maths question, and "what's the capital of Oman?".
  2. Print result.last_agent.name each time.
  3. Now delete both handoff_description values and re-run all three.
the third question belongs to neither specialist, and watching how your triage agent handles that is instructive. Without handoff descriptions, routing accuracy drops immediately — which tells you where to spend your effort.

Remembering the conversation

Each call to Runner.run is independent. The agent has no memory of the previous call unless you give it one, so the second question in a conversation needs the first question's context passed in somehow.

The manual way is explicit and sometimes exactly right:

manual_history.py
result = await Runner.run(agent, "What city is the Golden Gate Bridge in?")
new_input = result.to_input_list() + [
    {"role": "user", "content": "What state is it in?"}
]
result = await Runner.run(agent, new_input)
print(result.final_output)

to_input_list() returns the whole run — the user message, the model's messages, any tool calls and their results — as a list of input items. Append the next user message and pass the list as the input. You now own the history, which means you can trim it, store it, or redact it however you like.

The easier way is a session, a memory layer that loads history before each run and saves the new items afterwards:

session_chat.py
import asyncio
from agents import Agent, Runner, SQLiteSession

agent = Agent(name="Assistant", instructions="You are a helpful assistant")

async def main():
    session = SQLiteSession("user_123", "conversations.db")
    r1 = await Runner.run(agent, "What city is the Golden Gate Bridge in?", session=session)
    print(r1.final_output)
    r2 = await Runner.run(agent, "What state is it in?", session=session)
    print(r2.final_output)

asyncio.run(main())

The second question works because the session replayed the first exchange into the model's input. SQLiteSession("conversation_123") with no filename keeps everything in memory for the lifetime of the process; adding a filename makes it a real file you can inspect with the sqlite3 CLI, which is genuinely useful when debugging "why did the agent think that?".

The session ID is the conversation key, and in any multi-user application it must be derived from your authenticated user — one session per user or per thread. Note carefully that a session ID is not authentication: if you take it from a request body and trust it, any caller can read any conversation. Authorise first, then pick the session ID from the authorised identity.

SQLiteSession is for development. The SDK ships other backends for production — Redis, SQLAlchemy over Postgres, MongoDB, and OpenAI's own server-side Conversations API — and swapping between them is a one-line change because they all implement the same protocol. Which to choose, and the data-residency questions that come with it (a Gulf or Egyptian employer with a local-storage requirement will want the conversation history in a Postgres they control rather than in a vendor's conversation store), is a Mid-level topic.

Pick one history strategy per call The SDK also supports server-managed conversation state through conversation_id and previous_response_id. Combining either of those with a session= raises UserError: Session persistence cannot be combined with conversation_id, previous_response_id, or auto_previous_response_id. Choose local sessions or server-managed state, not both.
Try it
  1. Run the session script twice and notice that the second run already knows the earlier conversation, because it is on disk.
  2. Open conversations.db with sqlite3 conversations.db ".tables" and look at what was stored.
  3. Call await session.pop_item() between two runs and see the effect.
seeing the raw stored items demystifies memory entirely: it is a list of messages, loaded before the run and appended after it. Nothing more.

Streaming, and the throwaway REPL

Users will tolerate a slow answer that is visibly arriving and will not tolerate a spinner. Streaming is how you give them the first.

stream_tokens.py
import asyncio
from openai.types.responses import ResponseTextDeltaEvent
from agents import Agent, Runner

agent = Agent(name="Joker", instructions="You are a joker.")

async def main():
    result = Runner.run_streamed(agent, input="Please tell me 5 jokes.")
    async for event in result.stream_events():
        if event.type == "raw_response_event" and isinstance(event.data, ResponseTextDeltaEvent):
            print(event.data.delta, end="", flush=True)

asyncio.run(main())

The first thing to notice is that Runner.run_streamed is not awaited. It returns a result object immediately, and the work happens as you iterate stream_events(). Putting an await in front of it is a common early mistake.

There are three event types. raw_response_event carries the low-level model events, including the text deltas you just printed. run_item_stream_event carries higher-level happenings — a message was produced, a tool was called, a tool returned, a handoff was requested — and its name field takes one of a fixed set of values including message_output_created, tool_called, tool_output, handoff_requested and, spelled exactly like this, handoff_occured. That misspelling is deliberate and preserved for backward compatibility; do not "fix" it in your code, because the fixed spelling does not exist. The third type, agent_updated_stream_event, fires when a handoff changes the current agent, which is how you show a user "transferring you to billing".

The rule that catches people out: drain stream_events() to the end. The last token is not the end of the run. Session writes, approval bookkeeping and history compaction all happen after the final token, so if you break out of the loop as soon as you have the text you want, you can silently lose the conversation history you thought you were saving. If a stream appears to hang for a second or two after the last word, that is this post-processing, not a bug.

For exploration there is something even easier. run_demo_loop gives you an interactive terminal chat with your agent in one line:

repl.py
import asyncio
from agents import Agent, run_demo_loop

agent = Agent(name="Assistant", instructions="You are a helpful assistant")

asyncio.run(run_demo_loop(agent))

It streams by default, keeps the conversation going across turns, and exits when you type quit or exit or press Ctrl-D. This is the fastest way to get a feel for whether your instructions and tools actually work, and it is much more informative than re-running a script with one hard-coded question.

Try it
  1. Run the streaming script and watch the tokens arrive.
  2. Add an elif event.type == "run_item_stream_event": print(event.name) branch and run it again against a tool-using agent.
  3. Then run run_demo_loop on your weather agent and have an actual conversation with it.
the item-event names are a live narration of the agent loop you traced on paper earlier. Seeing tool_called then tool_output then message_output_created scroll past is the moment the loop stops being abstract.

Guardrails: your first safety net

Some requests should not reach your agent at all, and some answers should not reach your user. A guardrail is a function that inspects the input or the output and can stop the run by returning a tripwire.

guardrail.py
from pydantic import BaseModel
from agents import (
    Agent, GuardrailFunctionOutput, InputGuardrailTripwireTriggered,
    RunContextWrapper, Runner, TResponseInputItem,
)
from agents.decorators import input_guardrail

class HomeworkCheck(BaseModel):
    is_math_homework: bool

guardrail_agent = Agent(
    name="Homework check",
    instructions="Decide whether the user is asking you to do their maths homework.",
    output_type=HomeworkCheck,
)

@input_guardrail
async def math_guardrail(
    ctx: RunContextWrapper[None],
    agent: Agent,
    input: str | list[TResponseInputItem],
) -> GuardrailFunctionOutput:
    result = await Runner.run(guardrail_agent, input, context=ctx.context)
    return GuardrailFunctionOutput(
        output_info=result.final_output,
        tripwire_triggered=result.final_output.is_math_homework,
    )

support = Agent(
    name="Support",
    instructions="You help with study-skills questions.",
    input_guardrails=[math_guardrail],
)

Then you catch the tripwire where you call the agent:

guardrail_call.py
try:
    result = await Runner.run(support, "solve 2x+3=11")
    print(result.final_output)
except InputGuardrailTripwireTriggered as e:
    print("Refused:", e.guardrail_result.output.output_info)

A guardrail returns GuardrailFunctionOutput(output_info=..., tripwire_triggered=bool). The output_info is whatever you want to record about the decision — it ends up on the exception and in the trace, so put the reason in there. When tripwire_triggered is True, the run stops with InputGuardrailTripwireTriggered (or the output equivalent), and it is your job to catch that and say something sensible to the user.

Two facts about when guardrails run will save you real debugging time. Input guardrails run only on the first agent of a run — not on each agent after a handoff. Output guardrails run only on the agent that produces the final output. So in the triage example, an input guardrail on the triage agent protects the whole run, and an output guardrail belongs on each specialist that can finish.

And one about cost. By default input guardrails run in parallel with the agent, which is good for latency but means that by the time the tripwire fires, your agent may already have spent tokens and possibly run a tool. If the cost or the side effect matters, decorate with @input_guardrail(run_in_parallel=False) and the guardrail runs first and blocks. A cheap fast model for the guardrail agent, in blocking mode, is the standard shape for a moderation check.

Output guardrails have no parallel option — they cannot, since they need the output — and their signature is (ctx, agent, output).

Guardrails are not authorisation A guardrail is a filter on content. It is not a permission system. Tool arguments come from the model and are untrusted input, so the real authorisation check belongs inside the tool, where you can look at the authenticated user in ctx.context and refuse. Hiding a tool with is_enabled changes what the model sees, not what your code is allowed to do.
Try it
  1. Add the guardrail and run both an allowed and a blocked question.
  2. Print e.guardrail_result.output.output_info in the handler so you can see the reason.
  3. Switch to @input_guardrail(run_in_parallel=False) and compare the latency of a blocked request.
in blocking mode a refused request is faster and cheaper, because the main agent never ran. In parallel mode an allowed request is faster. That trade-off is the whole decision.

Choosing a model, and the settings that matter

If you do not specify a model, the SDK uses its default. As of version 0.20.0 that default is gpt-5.6-luna, which replaced gpt-5.4-mini; implicit defaults of reasoning.effort="none" and verbosity="low" go with it. Any guide you read that names an older default is out of date.

You can set the model in three places, and they have a strict precedence. From highest to lowest: an explicit model on the agent, then RunConfig.model for a single run, then the OPENAI_DEFAULT_MODEL environment variable, then the SDK default.

models.py
from agents import Agent, RunConfig, Runner

agent = Agent(name="A", model="gpt-5.6-sol")                 # per agent, highest priority

# or per run:
result = await Runner.run(agent, "Hi", run_config=RunConfig(model="gpt-5.6-sol"))
BASH
export OPENAI_DEFAULT_MODEL=gpt-5.6-sol

The habit to adopt now, before it costs you anything, is to set the model explicitly in anything you deploy. The SDK default has changed twice in recent minor versions. If you rely on it, an unrelated dependency bump can silently change which model serves your users, along with your latency and your bill. Explicit beats implicit here for the same reason you pin dependency versions.

Per-run and per-agent LLM parameters live in ModelSettings. You will not need most of its fields on day one, but four of them come up constantly:

settings.py
from agents import Agent, ModelSettings

agent = Agent(
    name="Summariser",
    instructions="Summarise the input in three bullet points.",
    model="gpt-5.6-luna",
    model_settings=ModelSettings(
        temperature=0.2,
        max_tokens=500,
        timeout=30.0,
        tool_choice="auto",
    ),
)

temperature you already know. max_tokens caps the output. timeout is in seconds and limits each model attempt, raising ModelTimeoutError if it is exceeded — worth setting, because the default is no timeout at all and a hung model call will hang your request. tool_choice controls whether the model may call tools: "auto" is the default, "required" forces it to call something, "none" forbids tools, and a tool's name forces that specific tool.

tool_choice="required" deserves a warning. Forcing a tool call on every turn is an easy way to build an infinite loop, since the model must call a tool, which produces a result, which triggers another turn, which must call a tool. The SDK protects you with reset_tool_choice, which defaults to True and clears the forced choice after a tool runs. If you ever turn that off, max_turns becomes your only backstop.

One note on retries, because it surprises people: the SDK does not retry model calls unless you configure ModelSettings.retry. There is no hidden layer absorbing transient 429s for you.

Try it
  1. Run the same prompt with temperature=0.0 and temperature=1.2, three times each.
  2. Set max_tokens=20 and see what a truncated response does to your program.
  3. Set timeout=0.001 and read the ModelTimeoutError.
step two matters more than it looks: on a non-streaming Responses call, a terminal status of failed or incomplete raises ModelBehaviorError as of version 0.22.0. Truncation is now an exception, not a quietly short answer.

Reading the errors you will actually hit

Every exception the SDK raises inherits from AgentsException, so you can catch that as a backstop. But the specific ones carry the information, and learning to read five of them makes you dramatically faster at this.

Error What really happened What to do
MaxTurnsExceeded: Max turns (10) exceeded The loop needed more than ten model calls, usually a tool-call loop Raise max_turns, fix the instructions, or handle it with a fallback
AgentRunner.run_sync() cannot be called when an event loop is already running. run_sync inside Jupyter, FastAPI or any async def Use await Runner.run(...)
InputGuardrailTripwireTriggered: Guardrail <Name> triggered tripwire A guardrail returned tripwire_triggered=True Catch it and answer the user; inspect e.guardrail_result
ModelBehaviorError: Invalid JSON input for tool <name> The model sent malformed tool arguments Simplify the tool schema, keep strict_mode=True, improve the docstring
UserError: The root of a strict JSON schema must be a non-nullable object... An output_type or tool schema strict mode cannot express Wrap it in a Pydantic model

MaxTurnsExceeded is the one you will see most, and it is almost always a symptom rather than a cause. Ten model calls for one question means the agent is going in circles: a tool that returns something the model does not understand, so it calls it again; instructions that tell it to verify everything twice; or tool_choice="required" with the loop protection disabled. Before you raise the limit, look at the trace and count what it actually did. If you genuinely need a long-running agent, you can set max_turns higher or to None, and you can install a graceful fallback:

error_handler.py
from agents import RunErrorHandlerResult, Runner

result = Runner.run_sync(
    agent,
    "Research the entire history of shipping.",
    max_turns=3,
    error_handlers={
        "max_turns": lambda d: RunErrorHandlerResult(
            final_output="That's too broad for me — can you narrow it down?",
            include_in_history=False,
        ),
    },
)

Two behaviours changed recently enough that older code gets them wrong. Since version 0.15.0, a model refusal raises ModelRefusalError with the refusal text on .refusal; before that it came back as an empty final_output or looped until the turn limit, so code that checks if not result.final_output to detect refusals is now checking the wrong thing. And as noted above, since 0.22.0 a non-streaming Responses call that ends failed or incomplete raises ModelBehaviorError rather than returning a partial result.

Finally, when you are stuck and the exception is not telling you enough, turn on verbose logging:

debug.py
from agents import enable_verbose_stdout_logging

enable_verbose_stdout_logging()

That prints what the SDK is sending and receiving, through the openai.agents logger. Note that model and tool payloads are redacted by default — controlled by the OPENAI_AGENTS_DONT_LOG_MODEL_DATA and OPENAI_AGENTS_DONT_LOG_TOOL_DATA environment variables, which both default to redacting. Setting one to 0 shows you the actual arguments the model invented, which is often exactly what you need when a tool keeps receiving nonsense. Do that in development only; those payloads contain user data.

Try it
  1. Write a tool that always returns "I don't know, ask the tool again" and an agent instructed to keep trying until it knows. Run it and read MaxTurnsExceeded.
  2. Add the max_turns error handler and run it again.
  3. Turn on verbose logging and re-run, with OPENAI_AGENTS_DONT_LOG_TOOL_DATA=0 set.
deliberately building the loop, then watching the loop in the logs, teaches more about the runner than any amount of reading. You will also see why the default redaction exists the moment your own prompts scroll past.

Seeing what happened: tracing

An agent run is a sequence of decisions you did not make, which makes "why did it do that?" the central debugging question. The SDK answers it with tracing, and the good news is that tracing is on by default: every run you have done in this guide has already produced a trace at https://platform.openai.com/traces.

The vocabulary is small. A trace is one end-to-end workflow, named "Agent workflow" unless you say otherwise. A span is one step inside it, and the SDK emits spans for the agent, each model generation, each function call, each guardrail and each handoff. Open a trace and you see the tree: this agent ran, it called the model, the model called get_weather with these arguments, the tool returned this, the model answered. Nearly every "why did it do that?" is answered by looking at that tree for sixty seconds.

Two things are worth setting even at beginner level:

tracing.py
from agents import RunConfig, Runner

result = await Runner.run(
    agent,
    "Where is my order?",
    run_config=RunConfig(
        workflow_name="Support triage",
        group_id=thread_id,
    ),
)

workflow_name means your traces are findable by name instead of being a hundred identical entries called "Agent workflow". group_id ties the separate runs of one conversation together, so a chat thread ID is the natural value.

You can turn tracing off globally with the OPENAI_AGENTS_DISABLE_TRACING=1 environment variable, or per run with RunConfig(tracing_disabled=True). There are two situations where you must think about this. If your organisation is on Zero Data Retention, OpenAI tracing is not available to you at all, and you need a third-party processor instead — the SDK supports a long list of them, including Langfuse, MLflow, Arize Phoenix and LangSmith, all of which have their own guides or near-equivalents in this catalogue. And if your data cannot leave a particular jurisdiction — a real constraint for banks and government work in the Gulf and in Egypt — then sending prompt contents to a tracing dashboard is a decision to make deliberately, not a default to inherit. RunConfig(trace_include_sensitive_data=False) keeps the shape of the trace while leaving the contents out.

Try it
  1. Run your tool-using agent, then open the Traces dashboard and find the run.
  2. Expand the function span and read the exact arguments the model generated.
  3. Re-run with workflow_name="my-first-agent" and group_id="test-1", and find it by name.
reading the real tool arguments is the single highest-value debugging habit in this guide. Most "the agent is broken" reports turn out to be "the model passed a value my tool did not expect", and the trace shows you that in one glance.

Putting it all together

Here is one small end-to-end project that uses everything above: a support assistant with a tool, structured output, a handoff, a session and a guardrail. It is short enough to read in one sitting and complete enough to be a template.

support_agent.py
import asyncio
from dataclasses import dataclass

from pydantic import BaseModel
from agents import (
    Agent, GuardrailFunctionOutput, InputGuardrailTripwireTriggered,
    ModelSettings, RunConfig, RunContextWrapper, Runner, SQLiteSession,
    TResponseInputItem, handoff,
)
from agents.decorators import input_guardrail, tool
from agents.extensions.handoff_prompt import RECOMMENDED_PROMPT_PREFIX


@dataclass
class AppContext:
    """Never sent to the model. Holds who is asking and how to look things up."""
    user_id: str
    orders: dict[str, str]


class Reply(BaseModel):
    answer: str
    needs_human: bool


@tool
def get_order_status(ctx: RunContextWrapper[AppContext], order_id: str) -> str:
    """Look up the delivery status of one order by its ID."""
    orders = ctx.context.orders
    if order_id not in orders:
        return f"No order {order_id} found for this account."
    return f"Order {order_id}: {orders[order_id]}"


@input_guardrail(run_in_parallel=False)
async def no_secrets(
    ctx: RunContextWrapper[AppContext],
    agent: Agent,
    input: str | list[TResponseInputItem],
) -> GuardrailFunctionOutput:
    text = input if isinstance(input, str) else str(input)
    leaked = "sk-" in text
    return GuardrailFunctionOutput(
        output_info="API key in message" if leaked else "clean",
        tripwire_triggered=leaked,
    )


refunds = Agent[AppContext](
    name="Refunds Agent",
    handoff_description="Handles refund requests and money back questions",
    instructions=f"{RECOMMENDED_PROMPT_PREFIX}\nExplain the refund policy and next steps. Never promise a date.",
    output_type=Reply,
    model="gpt-5.6-luna",
)

triage = Agent[AppContext](
    name="Support Agent",
    instructions=(
        f"{RECOMMENDED_PROMPT_PREFIX}\n"
        "You are a support agent for an online store. "
        "Use get_order_status before stating anything about a delivery. "
        "Hand off refund requests to the Refunds Agent. "
        "Set needs_human when you cannot resolve the question."
    ),
    tools=[get_order_status],
    handoffs=[handoff(refunds)],
    input_guardrails=[no_secrets],
    output_type=Reply,
    model="gpt-5.6-luna",
    model_settings=ModelSettings(temperature=0.2, timeout=30.0),
)


async def main() -> None:
    ctx = AppContext(user_id="u-42", orders={"A100": "out for delivery in Riyadh"})
    session = SQLiteSession(f"support:{ctx.user_id}", "support.db")
    config = RunConfig(workflow_name="Support triage", group_id=ctx.user_id)

    for question in [
        "Where is order A100?",
        "It arrived damaged, I want my money back.",
    ]:
        try:
            result = await Runner.run(
                triage, question, context=ctx, session=session, run_config=config
            )
            reply: Reply = result.final_output
            print(f"[{result.last_agent.name}] {reply.answer}")
            if reply.needs_human:
                print("  -> escalating to a human")
            print("  tokens:", result.context_wrapper.usage.total_tokens)
        except InputGuardrailTripwireTriggered as exc:
            print("Refused:", exc.guardrail_result.output.output_info)


asyncio.run(main())

Walk through what each piece earns. AppContext carries the user ID and the order lookup; because context is never sent to the model, the model cannot see other users' orders even if it asks. The tool takes ctx as its first parameter and reads from that context, so it can only look up orders for the authenticated account — authorisation inside the tool, exactly as the warning earlier insisted. Reply makes the output a typed object with an escalation flag, so your calling code branches on a boolean rather than searching prose for the phrase "speak to a human". output_type is set on both agents, because either can finish the run. The guardrail runs in blocking mode because the point is to stop before spending anything. The session key is derived from the authenticated user, not from user input. And RunConfig names the workflow and groups the runs so the traces are navigable.

Try it
  1. Run the project. Confirm that the first question is answered by the Support Agent and the second by the Refunds Agent.
  2. Send a third question containing the text sk-abc123 and watch the guardrail refuse it.
  3. Read the printed token count for each run, then add a third tool and a longer brief and compare it.
the token count is the number that makes agents feel real. A two-turn question costs a few thousand tokens; the same question with a chatty brief and five tools can cost ten times that, and now you can see it.

What you can now do, and what comes next

You can define an agent with instructions, tools and a model; run it synchronously, asynchronously or as a stream; make it return a validated Pydantic object; route between specialists with handoffs and know when to use as_tool instead; keep a conversation with a session; block requests with a guardrail; set the model and its settings explicitly rather than inheriting defaults; read the five errors that account for most beginner pain; and open a trace to find out what your agent actually did. That is a working, debuggable agent, and it is genuinely most of what day-to-day SDK work consists of.

Three habits are worth carrying forward from here, because they stay correct at every level. Set the model explicitly, so a dependency bump cannot change your product. Pin the minor version with openai-agents~=0.22.3, because minor releases may break public APIs. Read the trace before changing the prompt, because the trace usually tells you that the problem was a tool schema, not persuasion.

The Mid-level guide picks up where the defaults stop being good enough: how the loop behaves under load, tool_use_behavior and forcing tool choice, dynamic instructions, the full RunConfig surface, MCP servers as a tool source, production session backends, lifecycle hooks, retry and timeout policies, non-OpenAI models through LiteLLM, and deterministic testing with agents.testing.ScriptedModel. The Senior guide covers human-in-the-loop approvals and serialising RunState, the security model around that state, multi-tenancy, cost control, sandbox agents, and upgrade strategy across breaking minors.

Sideways, the natural neighbours in this catalogue are LangGraph if you want explicit graph control over the loop instead of letting the model decide, MCP for the protocol behind tool servers you did not write, Langfuse for tracing you host yourself, RAGAS for evaluating whether your agent is actually any good, and FastAPI for the service you will almost certainly wrap it in. The OpenAI API guide covers the Responses API underneath all of this, which is useful the first time you need to understand an error the SDK is merely relaying.

Sources