Skip to content
Back to student guides
AutoGenLLMsFrameworks & agents3 levels94 sectionsCovers AutoGen 0.7.5

The Complete AutoGen Guide

Build multi-agent conversations and event-driven agent systems with Microsoft AutoGen. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
15sections
16examples

This is part one of three. It covers everything you need to build working multi-agent programs with Microsoft AutoGen, starting from a machine that has never had it installed. By the end you can create an agent backed by a language model, give it Python functions to call, put two or three agents in a team that stops when it should, stream the conversation to your terminal, save and restore a run, and read the error messages AutoGen produces when something is wrong.

One thing has to be said before the first command, because it changes how you should read the rest: AutoGen is in maintenance mode. The README on github.com/microsoft/autogen states that it "will not receive new features or enhancements and is community managed going forward", and that "new users should start with Microsoft Agent Framework". AutoGen 0.7.5 is still a working, widely deployed framework, and the ideas in it — agents, teams, termination conditions, tool loops, handoffs — are the ideas the successor uses too. So learning it is not wasted. But you should know from the first page that this is a tool you learn to maintain and migrate, not one you bet a brand-new three-year project on. The second section covers that in detail, and every later section points out where the successor differs.

Each section ends with a Try it task. Do them as you go. Multi-agent code is deceptively easy to read and surprisingly easy to get wrong, and the only reliable cure is to watch your own run stop for the wrong reason once.

0.7.5latest release of autogen-agentchat, autogen-core, autogen-ext
Python 3.10+minimum supported interpreter
3 packagesagentchat (high level), core (runtime), ext (integrations)
Maintenancebug and security fixes only; successor is Microsoft Agent Framework

What AutoGen is, and the problem it solves

AutoGen is a Python framework for building applications in which one or more agents — programs that wrap a language model, a set of tools and a bit of memory — work on a task, and where some of them talk to each other to get it done. You describe the agents and how they take turns; AutoGen runs the loop, routes the messages, calls the tools, and tells you why it stopped.

To see what it buys you, write the thing it replaces. Suppose you want a model to answer questions about a city's weather. Without a framework you call the chat completions endpoint, read the response, notice that the model asked to call a function, find that function in a dictionary of your own making, call it, append the result to the message list in exactly the shape the provider expects, call the endpoint again, and check whether the model is done. That is forty or fifty lines of plumbing, none of it your idea. Now add a second model that reviews the first one's answer: you also own a turn-taking loop, a rule for when the conversation is finished, a cap so a disagreement cannot run forever, and a way to see what happened when it goes wrong at three in the morning.

Every line in that paragraph is a thing AutoGen owns for you. The tool-calling round trip becomes tools=[get_weather]. The turn-taking loop becomes RoundRobinGroupChat([writer, critic]). The rule for finishing becomes TextMentionTermination("APPROVE") | MaxMessageTermination(10). The visibility becomes await Console(team.run_stream(task=...)).

TASKa string you pass in
→
TEAMdecides who speaks
→
AGENTmodel + tools
→
TERMINATIONdecides when to stop
→
TASKRESULTmessages + stop_reason

That diagram is the whole shape of an AgentChat program, and almost everything in this guide is an elaboration of one of those five boxes.

It is worth knowing what came before, because you will meet the ancestor in search results. AutoGen started in 2023 as a research project whose central object was ConversableAgent and whose central verb was initiate_chat. That API — usually called v0.2 — was quick to demo and hard to reason about: agents replied through registered reply functions, the control flow lived inside the library, and configuration arrived as an llm_config dictionary. Microsoft rewrote it in early 2025. The v0.4 line, which 0.7.5 belongs to, split the framework into an event-driven runtime and a high-level task API, made every agent explicitly stateful, made asynchrony first class, and made the stopping rule a real object you can combine with | and &.

Try it
  1. Write out, in plain language, a task you would like two language models to do together — one producing something, one checking it.
  2. Underneath, write the rule that tells you the task is finished. Be exact: a word one of them says, a number of turns, a time limit.
  3. Keep the note. It becomes your termination condition later in this guide.
most people find the second part harder than the first. That asymmetry is why termination gets a whole section here: deciding what "done" means is the real design work in a multi-agent program.

Status first: maintenance mode, and what it means for your code

Treat this as part of the installation, not as a footnote. A framework's support status changes which habits are correct.

The facts, as of this guide's research date. The maintenance-mode banner went on the repository README on 2 October 2025 and was refreshed in April 2026. The last release published to PyPI is 0.7.5, dated 30 September 2025; commits after it are bug fixes, security fixes and documentation. The repository is not archived, so issues and pull requests still move, but no new capability is coming. The named successor is Microsoft Agent Framework, built by the AutoGen and Semantic Kernel teams together, which is at 1.x (the Python agent-framework package was at 1.19.0 in September 2026). Microsoft publishes an official AutoGen-to-Agent-Framework migration guide on Microsoft Learn.

Three practical consequences follow, and they shape the advice in every later section.

Pin everything, exactly. A maintained library absorbs the churn of its dependencies for you. A library in maintenance mode does not. The clearest live example is the Model Context Protocol extra: autogen-ext[mcp]==0.7.5 declares mcp>=1.11.0 with no upper bound, so a fresh install today pulls mcp 2.x, and the first import fails with ImportError: cannot import name 'RequestContext' from 'mcp.shared.context'. Nothing is wrong with your code; a dependency moved and nobody updated the bound. You fix it by pinning mcp<2 yourself. Assume more of these over time, and use a lock file from your very first project.

Read the release notes before any upgrade, even a minor one. AutoGen's 0.x minors carried real breaking changes — 0.4 to 0.5 to 0.6 to 0.7 each moved something. Pin autogen-agentchat, autogen-core and autogen-ext to the same exact version, which is what their own metadata expects: autogen-agentchat pins autogen-core to an exact match.

Learn the concepts with an eye on the mapping. The migration guide gives a direct table, and it is short enough to internalise now. AssistantAgent becomes Agent. OpenAIChatCompletionClient keeps its name. FunctionTool becomes a @tool decorator. AgentTool(agent) becomes agent.as_tool(). RoundRobinGroupChat becomes SequentialBuilder. MagenticOneGroupChat becomes MagenticBuilder. GraphFlow becomes WorkflowBuilder. Model context and state become AgentSession. One behavioural difference is worth memorising because it will bite you in both directions: AutoGen's AssistantAgent is effectively single-turn unless you raise max_tool_iterations, whereas the successor's Agent is multi-turn by default.

Do not start a greenfield production project on AutoGen If you are choosing a framework today for something that must still be running in 2028, Microsoft's own recommendation is Agent Framework. Learn AutoGen because you have inherited it, because your employer runs it, because it appears in interviews, and because the concepts transfer — not because it is the current best choice. Neighbouring options in this catalogue worth comparing are CrewAI, LangGraph, the OpenAI Agents SDK and Semantic Kernel.
Try it
  1. Open the repository README at github.com/microsoft/autogen and find the maintenance notice yourself.
  2. Open the releases page and read the notes for the most recent release.
  3. Write down one thing in those notes you do not understand yet, and look for it later in this guide.
reading a project's own release notes before writing a line against it is the single habit that separates engineers who get surprised from engineers who do not. Start it here, where the notes are short.

Three packages and one very confusing name

AutoGen ships as three Python packages, and understanding the split tells you where to look for anything.

Package Import name What lives in it
autogen-core autogen_core The event-driven runtime, the base types, message passing, agent identity, tool and memory interfaces
autogen-agentchat autogen_agentchat The high-level, task-driven API: preset agents, teams, termination conditions, the Console UI
autogen-ext autogen_ext Everything that touches the outside world: model clients, code executors, memory backends, MCP, extra runtimes

Beginners should live almost entirely in AgentChat, reaching into ext for a model client and into core for a couple of helper types. Core's own API — runtimes, topics, subscriptions, routed agents — is a different and lower-level way to write the same kinds of programs, and it belongs at mid and senior level. The three-layer split exists so that AgentChat can stay opinionated while Core stays general.

Alongside the libraries are three developer tools you will see mentioned, none of them needed to learn the framework. AutoGen Studio is a low-code browser UI for assembling teams by clicking; its own documentation says it is "not meant to be a production-ready app", and the stable release 0.4.2.2 requires the autogen-* packages below 0.6, so it must live in a separate virtual environment from your 0.7.x code. Magentic-One is a prebuilt generalist team plus an m1 command line, installed from magentic-one-cli. AgBench is a benchmarking harness.

Now the name problem, which costs beginners more time than any technical issue in this guide.

Microsoft AutoGen — what this guide teaches

  • Install: autogen-agentchat, autogen-core, autogen-ext
  • Imports use underscores: from autogen_agentchat.agents import AssistantAgent
  • Docs: microsoft.github.io/autogen/stable/
  • Objects: AssistantAgent, RoundRobinGroupChat, await agent.run(...)

Not this guide

  • PyPI autogen and ag2 belong to AG2, the community fork at ag2ai/ag2
  • import autogen is always v0.2-style or AG2 code
  • Docs: docs.ag2.ai — a different, v0.2-descended API
  • Objects: ConversableAgent, initiate_chat, llm_config

There is a third name in the mess. pyautogen was the original distribution, and the migration guide states that Microsoft "no longer [has] admin access to the pyautogen PyPI package, and the releases from that package are no longer from Microsoft since version 0.2.34". Today pyautogen 0.10.0 is a thin proxy depending on autogen-agentchat, so installing it is harmless but pointless. The honest rule is the import rule: if the code says import autogen, it is not the framework you installed.

A fast triage for any tutorial you find Search the page for ConversableAgent, initiate_chat, register_reply, llm_config, config_list or cache_seed. A hit on any of them means the page predates the rewrite or belongs to the fork, and you should close it. A page that says AssistantAgent, model_client= and await is written for what you have.
Try it
  1. Search the web for "autogen tutorial" and open the first three results.
  2. Apply the triage above to each one and label it current, v0.2 or AG2.
  3. Note how many of the three you would have followed blindly.
usually at least one of the three is unusable. Doing this consciously once makes the check automatic, and it saves the classic lost afternoon chasing an AttributeError on ConversableAgent.

The mental model: four nouns

Everything in AgentChat is built from four nouns. Learn them in this order, because each one depends on the ones before it.

The model client is the thing that talks to a language model. Its interface is ChatCompletionClient, defined in autogen_core.models, and the concrete implementations live in autogen_ext.models.* — one for OpenAI and Azure OpenAI, one for Anthropic, one for Ollama, one for llama.cpp, one for Azure AI, an adapter for Semantic Kernel, and a ReplayChatCompletionClient that returns canned strings and is perfect for tests. A client exposes create() and create_stream() for calls, count_tokens() and remaining_tokens() for budgeting, actual_usage() and total_usage() for cost, a model_info property describing what the model can do, and close(). You create one client and share it between agents.

The agent has a name that must be a valid Python identifier, a description that teams use to decide who should speak, and the behaviour you configure: a system message, tools, memory, a view of the conversation history. The important property, and the one beginners most often get wrong, is that agents are stateful. An agent remembers the conversation across calls. So when you call it again you pass only the new message, never the whole transcript — passing the transcript duplicates everything in its context and doubles your token bill for no benefit.

The team is a group of agents sharing one message thread, driven by a group-chat manager. Four presets matter at this level, and RoundRobinGroupChat is the only one you need today. It gives agents a fixed turn order. SelectorGroupChat asks a model to pick the next speaker each turn. Swarm lets agents hand off to each other explicitly. MagenticOneGroupChat wraps an orchestrator that plans, tracks whether progress has stalled, and re-plans. A fifth, GraphFlow, lets you draw the control flow as a directed graph and is still documented as experimental.

The termination condition decides when a run stops. It is a stateful callable, checked after each agent response against only the new messages, which returns a StopMessage or None. You combine conditions with | for "or" and & for "and", and each one resets itself after a run finishes.

Two return types carry results. Response, returned by an agent's low-level on_messages, has chat_message (the final message) and inner_messages (the events that led to it). TaskResult, returned by run() and yielded last by run_stream(), has messages and — the field you will read most often in your life — stop_reason, a human-readable string saying exactly why the loop ended.

Messages come in two families, and knowing which is which stops a lot of confusion when you start streaming. Chat messages travel between agents: TextMessage, MultiModalMessage, StopMessage, HandoffMessage, ToolCallSummaryMessage and the generic StructuredMessage[T]. Events are internal to one agent and exist so you can watch it think: ToolCallRequestEvent, ToolCallExecutionEvent, MemoryQueryEvent, UserInputRequestedEvent, ModelClientStreamingChunkEvent, ThoughtEvent and SelectSpeakerEvent. Every message carries source, models_usage, metadata, id and created_at, and offers to_text().

autogen_agentchat
TeamRoundRobin, Selector, Swarm
AgentAssistantAgent, UserProxyAgent
Terminationcombined with | and &
autogen_core
Runtimedelivers messages, owns lifecycles
InterfacesChatCompletionClient, Tool, Memory
autogen_ext
Model clientsOpenAI, Anthropic, Ollama
ExecutorsDocker, Jupyter
MemoryChroma, Redis, Mem0
Try it
  1. Say the four nouns out loud, each with one sentence of your own words.
  2. For your task from the first section, name the agents, say which team preset fits, and state the termination condition.
  3. Decide which model you will use, and whether it supports function calling.
a one-paragraph design. Everything from here is typing. The last question matters more than it looks: an agent with tools= on a model whose model_info says function_calling: False raises ValueError: The model does not support function calling. before it ever reaches the network.

Installing AutoGen and checking it works

Use a virtual environment. AutoGen has a wide dependency surface and version pinning matters more here than in a maintained library, so a per-project environment is not optional advice.

BASH
python3 -m venv .venv
source .venv/bin/activate              # Linux and macOS
# Windows cmd:        .venv\Scripts\activate.bat
# Windows PowerShell: .venv\Scripts\Activate.ps1

If you prefer conda, conda create -n autogen python=3.12 && conda activate autogen is equivalent. Python 3.10 is the minimum the packages declare; 3.12 is a safe choice.

Then install. The first line is the official README's; the second is what you should commit.

BASH
pip install -U "autogen-agentchat" "autogen-ext[openai]"

# For reproducibility, pin the exact versions instead:
pip install "autogen-agentchat==0.7.5" "autogen-ext[openai]==0.7.5"

autogen-core arrives automatically, because autogen-agentchat pins it to the same exact version. The square brackets are pip extras: autogen-ext is a bag of optional integrations and you install only the ones you need. The ones worth knowing early are openai, anthropic, ollama, gemini, azure and llama-cpp for models; mcp, docker and jupyter-executor for tools and code execution; redis, chromadb, mem0 and diskcache for memory and caching; magentic-one, web-surfer and file-surfer for the prebuilt agents; and grpc and rich for the distributed runtime and prettier terminal output. In PowerShell, keep the quotes around the whole argument or the shell eats the brackets.

Set your key as an environment variable. OpenAIChatCompletionClient reads OPENAI_API_KEY by itself, and the Anthropic client reads ANTHROPIC_API_KEY.

BASH
export OPENAI_API_KEY=sk-...           # PowerShell: $env:OPENAI_API_KEY="sk-..."

Now verify, in two steps. First that the packages are importable and the versions agree:

BASH
python -c "import autogen_agentchat, autogen_core; print(autogen_agentchat.__version__, autogen_core.__version__)"
# -> 0.7.5 0.7.5
pip show autogen-agentchat autogen-core autogen-ext

Second — and do this before you spend a single token — run the framework end to end with no model at all. ReplayChatCompletionClient is a real client that returns the strings you give it in order. It proves your installation, your async setup and your mental model are all correct, and it costs nothing.

smoke_test.py
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.replay import ReplayChatCompletionClient

async def main():
    agent = AssistantAgent("assistant", model_client=ReplayChatCompletionClient(["Hello World!"]))
    result = await agent.run(task="Say hello")
    print(result.messages[-1].content)   # Hello World!

asyncio.run(main())

Three per-platform notes save a lot of grief. On macOS and Linux use python3; Docker (Docker Desktop or Colima on macOS, Docker Engine on Linux) is needed for the Docker code executor, and autogen-ext[web-surfer] also needs playwright install. On Windows, the local code executor needs the Proactor event loop, and the source warns "The current event loop policy is not WindowsProactorEventLoopPolicy…"; fix it with asyncio.set_event_loop_policy(asyncio.WindowsProactorEventLoopPolicy()). In Jupyter or IPython do not call asyncio.run(...) — the notebook already runs an event loop, and you will get RuntimeError: asyncio.run() cannot be called from a running event loop. Use top-level await.

Try it
  1. Create the virtual environment, install the pinned versions, and run the two verification commands.
  2. Save and run smoke_test.py. Confirm it prints Hello World! with no API key set.
  3. Run pip freeze > requirements.txt and open the file.
a working environment and a lock of it. The third step is the one people skip; in a framework whose dependency bounds are no longer being updated it is the difference between an environment you can rebuild next month and one you cannot.

Your first real agent

Now the same program against a real model. This is the official hello world, and it is worth reading line by line because every later example is a variation of it.

hello.py
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient

async def main() -> None:
    model_client = OpenAIChatCompletionClient(model="gpt-4.1")
    agent = AssistantAgent("assistant", model_client=model_client)
    print(await agent.run(task="Say 'Hello World!'"))
    await model_client.close()

asyncio.run(main())

Four things are happening. The client is created with a model name and no key, because it falls back to OPENAI_API_KEY. The agent is created with a name — "assistant" is a valid Python identifier, which is required — and the client. run(task=...) sends one message and returns a TaskResult, which is why printing it gives you the whole structure rather than just the text. And close() releases the client's HTTP connections; forget it and Python will complain about unclosed sessions at exit.

AssistantAgent has more knobs than any other class you will touch, so here are the ones that matter at this level, with their real defaults.

Parameter Default What it does
system_message A built-in instruction ending "Reply with TERMINATE when the task has been completed." The persona and the rules. Change it for almost every real agent
description "An agent that provides assistance with ability to use tools." How a team decides this agent should speak. Worth writing properly once you have more than one agent
tools None Python callables the model may invoke
max_tool_iterations 1 How many model-then-tool rounds happen inside one turn
reflect_on_tool_use None Whether the model gets a chance to summarise the tool result in prose
model_client_stream False Whether token chunks are emitted as events
output_content_type None A Pydantic model for structured output
memory None Memory stores queried before each model call
model_context None (unbounded) Which slice of history is sent to the model

Note that default system message. The stock AssistantAgent has been told to say TERMINATE when it is finished, which is why so many examples pair it with TextMentionTermination("TERMINATE"). If you replace the system message — and you should, for any agent with a real job — you also take responsibility for whatever word your termination condition is looking for. A termination condition watching for a word your agent was never told to say produces a run that goes to its message cap every time, and the only clue is stop_reason.

There are three ways to run things, and they differ only in how much you see:

  • result = await agent.run(task="...") gives you the TaskResult at the end and nothing in between.
  • async for message in agent.run_stream(task="..."): ... yields each message and event as it happens, with the TaskResult last.
  • await Console(agent.run_stream(task="...")) prints that stream to your terminal in a readable form. Console comes from autogen_agentchat.ui, and Console(..., output_stats=True) adds token counts.

task can be a plain string, a BaseChatMessage, or a list of them. Both run and run_stream also take cancellation_token= to abort and output_task_messages= (default True) to control whether your own task is echoed back in the result.

Try it
  1. Run hello.py and read the printed TaskResult. Find stop_reason in it.
  2. Replace the print with await Console(agent.run_stream(task="Explain what a container is in three sentences.")) and import Console from autogen_agentchat.ui.
  3. Give the agent system_message="You answer only in bullet points, in at most 40 words." and run it again.
the same program with three levels of visibility, and proof that the system message is the cheapest control you have. Keep Console in your scratch scripts permanently; debugging an agent you cannot see is guesswork.

Giving an agent tools

A tool is how an agent does something other than produce text. In AutoGen, a tool is any Python function — synchronous or asynchronous — with type hints on its parameters and a docstring. You pass the function itself and AutoGen wraps it in a FunctionTool for you. The type hints become the schema the model sees, and the docstring becomes the description the model uses to decide whether to call it, so both are load-bearing rather than decorative.

weather.py
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient

async def get_weather(city: str) -> str:
    """Get the weather for a given city."""
    return f"The weather in {city} is 73 degrees and Sunny."

async def main() -> None:
    model_client = OpenAIChatCompletionClient(model="gpt-4.1")
    agent = AssistantAgent(
        "weather_agent",
        model_client=model_client,
        tools=[get_weather],
        reflect_on_tool_use=True,
    )
    await Console(agent.run_stream(task="What is the weather in Dubai?"))
    await model_client.close()

asyncio.run(main())

Run it with Console and read the stream carefully, because it shows the tool protocol in full: a ToolCallRequestEvent where the model asks for get_weather(city="Dubai"), a ToolCallExecutionEvent carrying what your function returned, and then a final message.

What that final message is depends on reflect_on_tool_use. With it set to True, the tool result goes back to the model and the agent's final message is the model's prose. With it left off, the agent's final message is a ToolCallSummaryMessage containing the raw tool result formatted with tool_call_summary_format, which defaults to "{result}". Neither is wrong. Reflection reads better and costs an extra model call; the summary is cheaper, faster and exactly what you want when the tool's output is already the answer. The default is subtle: reflect_on_tool_use=None resolves to True when output_content_type is set and False otherwise.

One tool round by default AssistantAgent.max_tool_iterations defaults to 1, which means a single model-then-tool round per turn. If your task needs the agent to call a tool, look at the result, and then call another tool, it will not do it — it will stop after the first. Raise the limit explicitly, for example max_tool_iterations=10. This parameter arrived in 0.6.2, so a great many tutorials written before it do not mention it, and their authors worked around the limit by wrapping the agent in a team.

Two rules about tools are worth learning now rather than discovering through an exception. First, an agent takes tools= or workbench=, never both; passing both raises ValueError: Tools cannot be used with a workbench. A workbench is a set of tools that share state and resources — StaticWorkbench for a fixed list, McpWorkbench for a Model Context Protocol server — and it is the mid-level topic. Second, the model must support function calling. If model_info says function_calling: False, passing tools= raises ValueError: The model does not support function calling. straight away.

Tools are also where security enters. A tool runs with your program's privileges, and its arguments come from a model that may be acting on text a stranger wrote. Validate them inside the function as you would an HTTP request body: check ranges, restrict paths, never interpolate a model-supplied string into a shell command or SQL statement. The framework does not do this for you, and the same caution applies with more force to MCP servers — the README's warning is "Only connect to trusted MCP servers." See the MCP guide.

Try it
  1. Run weather.py and identify the two tool events in the Console output.
  2. Remove reflect_on_tool_use=True and run it again. Compare the final messages.
  3. Add a second tool, def convert_f_to_c(f: float) -> float, and ask for the weather in Dubai in Celsius. Then set max_tool_iterations=5 and ask again.
the third step is the lesson. With one iteration the agent cannot chain two tools; with five it can. This single default explains a large share of "my agent just stops halfway" questions.

Two agents in a team, and knowing when to stop

A team is a group of agents sharing one message thread. The simplest and most predictable is RoundRobinGroupChat, which gives them a fixed turn order. The classic shape is a producer and a reviewer.

team.py
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient

async def main() -> None:
    model_client = OpenAIChatCompletionClient(model="gpt-4.1")

    primary = AssistantAgent(
        "primary",
        model_client=model_client,
        description="Writes the first draft.",
        system_message="You are a helpful writer. Write short, concrete prose.",
    )
    critic = AssistantAgent(
        "critic",
        model_client=model_client,
        description="Reviews drafts and approves them.",
        system_message=(
            "Review the draft and give specific, actionable feedback. "
            "When the draft needs no further change, reply with the single word APPROVE."
        ),
    )

    termination = TextMentionTermination("APPROVE") | MaxMessageTermination(10)
    team = RoundRobinGroupChat([primary, critic], termination_condition=termination)

    result = await Console(team.run_stream(task="Write a short poem about the Red Sea in autumn."))
    print(result.stop_reason)

    await team.reset()
    await model_client.close()

asyncio.run(main())

Read the termination condition as English: stop when someone says APPROVE, or when ten messages have gone by. The second half is not optional pessimism, it is the safety net. A team with no termination condition and no max_turns runs until your budget or the provider stops it, and two polite agents can disagree about a poem for a remarkably long time. Make the pairing a habit: one condition that expresses what you actually want, one that bounds the damage.

Here are the built-in conditions, all from autogen_agentchat.conditions:

Condition Stops when
MaxMessageTermination(max_messages, include_agent_event=False) A message count is reached
TextMentionTermination(text, sources=None) Some text appears in a message, optionally only from named sources
TextMessageTermination(source=None) A TextMessage is produced, optionally by a named source
StopMessageTermination() An agent produces a StopMessage
TokenUsageTermination(max_total_token=..., max_prompt_token=..., max_completion_token=...) A token budget is spent
TimeoutTermination(timeout_seconds) Wall-clock time runs out
HandoffTermination(target) A handoff to a given target happens, typically "user"
SourceMatchTermination(sources) One of the named agents has spoken
FunctionCallTermination(function_name) A particular tool is called
ExternalTermination() You call .set() on it from elsewhere
FunctionalTermination(func) Your own function says so

One counting detail trips up nearly everybody: MaxMessageTermination counts the task message. With MaxMessageTermination(3) you get your task plus two agent replies, and the stop reason reads "Maximum number of messages 3 reached, current message count: 3". If you want N replies, ask for N+1.

The other behaviour to internalise is run and resume. Calling team.run() again — with no task, or with a new one — continues from where the previous run stopped; the agents still remember everything. await team.reset() is what clears the state and gives you a clean slate. This is a feature, and it is also the reason a long-running script slowly gets more expensive if you never reset: the context keeps growing.

You will meet the other presets at mid level, but a one-line map helps now. SelectorGroupChat spends an extra model call each turn asking which agent should speak next, needs at least two participants, and by default will not let the same agent speak twice in a row. Swarm routes by explicit handoffs, where the next speaker is whoever the most recent HandoffMessage named. MagenticOneGroupChat runs an orchestrator with a task and progress ledger, defaults to twenty turns, and gives up after three stalls. Since 0.7.1 a team can itself be a participant of the first two, which is how you nest a sub-team.

Try it
  1. Run team.py and read the printed stop_reason.
  2. Delete | MaxMessageTermination(10), change APPROVE in the critic's system message to PERFECT, but leave the condition watching for APPROVE. Run it with a low max_turns=4 on the team so you can stop it safely.
  3. Restore the pair, then call team.run() a second time with task="Now make it two lines shorter." without resetting.
step two is a deliberate mismatch between what the agent was told to say and what the condition watches for, and it is the single most common termination bug. Step three shows resumption: the team already knows which poem you mean.

Seeing what happened: streaming, events and state

Debugging agents is mostly a visibility problem, so spend ten minutes getting good at this and you will save hours later.

run_stream yields every chat message and every event in order, with the TaskResult last. Wrapping it in Console is the quick option; iterating it yourself is how you build your own UI or log.

watch.py
from autogen_agentchat.base import TaskResult

async for item in team.run_stream(task="Write a haiku about Cairo traffic."):
    if isinstance(item, TaskResult):
        print("STOPPED:", item.stop_reason)
    else:
        print(f"[{item.source}] {type(item).__name__}: {item.to_text()[:120]}")

Setting model_client_stream=True on an AssistantAgent adds ModelClientStreamingChunkEvents, so you see tokens as they arrive rather than complete messages. That is what you want behind a chat interface, and it is noise in a log.

Token usage is attached to the messages themselves. Every message carries models_usage, a RequestUsage with prompt_tokens and completion_tokens, and the client keeps running totals you can read with total_usage() and actual_usage(). Console(..., output_stats=True) prints them for you. Look at these numbers on your very first team run; the shape of the bill is rarely what people guess, because each agent re-reads the whole shared thread on its turn.

State is the other half of visibility. save_state() on a team or an agent returns a JSON-serialisable dictionary, and load_state() restores it.

persist.py
import json

state = await team.save_state()
with open("state.json", "w") as f:
    json.dump(state, f)

# later, in another process
with open("state.json") as f:
    await new_team.load_state(json.load(f))

For an AssistantAgent, that state is essentially its model context — the conversation it remembers. This is how you build a web application on top of AgentChat: one team per user session, state written to Postgres or Redis between requests, loaded again on the next one. You cannot load_state while a team is running, and trying raises RuntimeError: The team cannot be loaded while it is running.

Two related constraints surprise people building their first service. A team instance cannot run twice concurrently: a second overlapping run() raises ValueError: The team is already running, it cannot run again until it is stopped. Never share one team object across HTTP requests. And for controlled stopping, prefer ExternalTermination().set(), which lets the current agent finish its turn, over a CancellationToken, which aborts immediately and can leave state inconsistent.

Finally, plain Python logging works: autogen_core exposes TRACE_LOGGER_NAME for human-readable debug output and EVENT_LOGGER_NAME for structured events. The runtime is also instrumented with OpenTelemetry, emitting GenAI spans named create_agent, invoke_agent and execute_tool, so any OTLP backend can show you a trace of a run — that is how you would wire it to Langfuse. AUTOGEN_DISABLE_RUNTIME_TRACING=true switches the runtime spans off.

Try it
  1. Rewrite your team run with the explicit async for loop above and watch the event types scroll past.
  2. Add output_stats=True to a Console call and note the total tokens for a two-agent, six-message run.
  3. Save the team state to state.json, open the file, then load it into a freshly constructed team and continue the conversation.
reading state.json by hand is the moment the framework stops being magic: it is a conversation, a few counters, and your configuration. Anything you can save you can inspect, diff and test.

Putting a human in the loop

Sometimes the right next speaker is a person. AutoGen gives you two ways to arrange that, and choosing wrongly is a classic first-project mistake.

The direct way is UserProxyAgent, an agent whose turn is taken by a human. Its signature is UserProxyAgent(name, *, description="A human user", input_func=None) and input_func defaults to Python's input().

human.py
from autogen_agentchat.agents import AssistantAgent, UserProxyAgent
from autogen_agentchat.conditions import TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat

assistant = AssistantAgent("assistant", model_client=model_client)
user = UserProxyAgent("user_proxy")

team = RoundRobinGroupChat(
    [assistant, user],
    termination_condition=TextMentionTermination("APPROVE", sources=["user_proxy"]),
)
await Console(team.run_stream(task="Draft a two-sentence summary of our data residency policy."))

This is perfect for a terminal script and wrong for a web application, for a reason worth understanding. While UserProxyAgent waits for input it blocks the running team, and a running team cannot be saved. So a web request that is waiting for a human is a request holding a live team object with unsaveable state, which is exactly the thing you cannot do at scale.

The pattern that does work is to make "we need the human" a reason to stop. End the run, return to your application, and call run() again with the person's reply when it arrives. HandoffTermination(target="user") and a plain TextMentionTermination both do this, and because resuming a team is normal, the second run() picks up precisely where the first stopped. You can watch for UserInputRequestedEvent in the stream to know a request for input is coming.

A rule of thumb for human input Synchronous human — someone at a terminal, now — use UserProxyAgent. Asynchronous human — someone who will reply within the next hour, through a web form or a chat app — terminate, persist the state, and resume. The second shape also survives a process restart, which the first does not.
Try it
  1. Run human.py and reply a couple of times, finishing with APPROVE.
  2. While it is waiting for your input, try to picture what await team.save_state() would have to capture. Then read the paragraph above again.
  3. Rewrite it without UserProxyAgent: terminate on a word, print the draft, use input() in your own code, and call team.run(task=your_reply).
two programs that behave identically at a terminal and differ completely in how they would deploy. The second is the one you would put behind an HTTP endpoint.

Context, memory and keeping the bill down

By default an agent sends its entire conversation history to the model on every turn. That is the UnboundedChatCompletionContext, and for a short task it is exactly right. For a long one it is a slow-motion cost problem: every turn is longer than the last, and eventually you hit the model's context limit and the run fails.

A model context is the configurable answer to "which part of history does the model actually see". There are four, all in autogen_core.model_context:

Model context Sends
UnboundedChatCompletionContext Everything. The default
BufferedChatCompletionContext(buffer_size=N) The last N messages
HeadAndTailChatCompletionContext(head_size=, tail_size=) The first few and the last few, dropping the middle
TokenLimitedChatCompletionContext(model_client, token_limit=) As much as fits in a token budget

HeadAndTailChatCompletionContext encodes a real insight: the beginning of a conversation holds the task and the constraints, the end holds the current state, and the middle is the most droppable part.

Memory is the complementary idea. A memory store is queried before each model call and its results are injected into the context, which lets an agent carry facts across conversations rather than within one. The interface is autogen_core.memory.Memory, and the simplest implementation is ListMemory.

memory.py
from autogen_core.memory import ListMemory, MemoryContent, MemoryMimeType
from autogen_core.model_context import BufferedChatCompletionContext

mem = ListMemory()
await mem.add(MemoryContent(content="User prefers metric units", mime_type=MemoryMimeType.TEXT))

agent = AssistantAgent(
    "assistant",
    model_client=model_client,
    model_context=BufferedChatCompletionContext(buffer_size=10),
    memory=[mem],
)

When memory fires you will see a MemoryQueryEvent in the stream, which is how you confirm it is doing anything. Beyond ListMemory, autogen_ext offers vector and service-backed stores — ChromaDBVectorMemory, RedisMemory, Mem0Memory and a canvas memory — each behind its own pip extra.

While we are on cost, collect the levers in one place, because this is the question your manager will ask. Cap the conversation with MaxMessageTermination or max_turns. Cap the spend directly with TokenUsageTermination. Cap the context with a buffered or token-limited model context. Avoid SelectorGroupChat when routing is deterministic, since it pays for an extra model call every turn just to choose a speaker. Leave reflect_on_tool_use off when the tool's raw output is already the answer, since reflection is another model call. And cache repeated calls: model-response caching has been off by default since v0.4, and you turn it on by wrapping your client.

cache.py
from autogen_ext.models.cache import ChatCompletionCache, CHAT_CACHE_VALUE_TYPE
from autogen_ext.cache_store.diskcache import DiskCacheStore
from diskcache import Cache

cached_client = ChatCompletionCache(
    model_client,
    DiskCacheStore[CHAT_CACHE_VALUE_TYPE](Cache("/tmp/autogen-cache")),
)

That needs pip install "autogen-ext[diskcache]", and there is a Redis equivalent in autogen_ext.cache_store.redis. While you are developing and running the same prompt twenty times, a cache is the single biggest saving available to you.

Try it
  1. Run a six-turn team conversation with Console(..., output_stats=True) and record the total tokens.
  2. Give both agents model_context=BufferedChatCompletionContext(buffer_size=4) and run the same task again. Compare.
  3. Add a ListMemory entry stating a preference, ask a question that should respect it, and find the MemoryQueryEvent in the stream.
a number you can show someone, plus the understanding that context is a choice. Step two also sometimes makes the answer worse, which is the real trade-off: you are deciding what the model is allowed to forget.

Configuration, and the errors you will actually see

Two configuration topics matter at this level: telling AutoGen about a model it does not recognise, and saving a setup as JSON.

The first is model_info, a small dictionary describing what a model can do. The OpenAI client knows OpenAI's own models, so for gpt-4.1 you pass nothing. The moment you point it at something else — a local model through Ollama's OpenAI-compatible endpoint, a vLLM or LM Studio server, a LiteLLM proxy, Gemini's OpenAI-compatible endpoint — it has no idea, and you must say.

local_model.py
from autogen_ext.models.openai import OpenAIChatCompletionClient

client = OpenAIChatCompletionClient(
    model="llama3.2",
    base_url="http://localhost:11434/v1",
    api_key="placeholder",
    model_info={
        "vision": False,
        "function_calling": True,
        "json_output": False,
        "family": "unknown",
        "structured_output": True,
    },
)

Omit it and you get ValueError: model_info is required when model name is not a valid OpenAI model. Omit just the structured_output key and you get UserWarning: Missing required field 'structured_output' in ModelInfo. This field will be required in a future version of AutoGen. — a warning today, an error later, so fill it in. Lie about function_calling and your tool calls fail at the provider instead of in AutoGen, which is a much worse place to debug. family takes a value from ModelFamily or the string "unknown". Note also that model_capabilities= is the deprecated predecessor of model_info and the two are mutually exclusive; if you see it in a tutorial, rename it.

A direct Ollama client also exists at autogen_ext.models.ollama.OllamaChatCompletionClient, taking a host key, which is tidier than the OpenAI-compatible endpoint. The Ollama guide covers the server side.

The second topic is component configuration. Every AutoGen component can serialise itself to JSON with dump_component() and be rebuilt with load_component().

declarative.py
cfg = team.dump_component()
json_str = cfg.model_dump_json()

team2 = RoundRobinGroupChat.load_component(cfg)

This is how AutoGen Studio stores teams, and it keeps configuration out of code. Two warnings come with it. Callables are not serialisable — a selector_func, a GraphFlow lambda condition and a FunctionTool cannot survive the round trip. And dump_component() writes an explicitly passed api_key into the JSON, so rely on environment variables and never commit a dumped config. There is a sharper point too: in released 0.7.5, load_component() imports whatever module path the provider string names, so loading component JSON from an untrusted source is remote code execution. A namespace restriction exists on main but is in no PyPI release. Treat component JSON as code.

Now the errors. AutoGen's messages are unusually good — most of them name the fix — so learning to read them is quick.

Message What it means
openai.OpenAIError: Missing credentials. Please pass an api_key… No key. export OPENAI_API_KEY=sk-... or pass api_key=
openai.AuthenticationError: Error code: 401 … 'code': 'invalid_api_key' The key is wrong or revoked
openai.APIConnectionError: Connection error. Wrong base_url, server down, or TLS interception by a corporate proxy. Check the URL and your CA bundle
ValueError: The agent name must be a valid Python identifier. You wrote "my agent" or "my-agent". Use my_agent
ValueError: The participant names must be unique. Two agents in one team share a name
ValueError: The model does not support function calling. tools= on a model whose model_info says otherwise
ValueError: Tools cannot be used with a workbench. You passed both. Pick one
ValueError: The team is already running, it cannot run again until it is stopped. Concurrent run() on one team object. One team per session
RuntimeError: The team cannot be loaded while it is running. load_state mid-run
ValueError: At least two participants are required for SelectorGroupChat. Use RoundRobinGroupChat for one agent
RuntimeError: asyncio.run() cannot be called from a running event loop asyncio.run inside Jupyter. Use top-level await
ModuleNotFoundError: No module named 'anthropic' (or ollama, chromadb, playwright) A missing pip extra. Install "autogen-ext[anthropic]" and so on
ImportError: cannot import name 'RequestContext' from 'mcp.shared.context' mcp 2.x with autogen-ext 0.7.5. Pin "mcp<2"
ModuleNotFoundError: No module named 'autogen' or AttributeError on ConversableAgent You are following a v0.2 or AG2 tutorial

Four non-error pitfalls round out the list, and each one produces behaviour rather than a traceback, which makes them harder. A team with no termination condition and no max_turns runs until something external stops it. Passing the full history into agent.run() each time duplicates context, because agents are already stateful. Forgetting await model_client.close() leaves HTTP connections open and produces warnings at exit. And max_tool_iterations=1 means one tool round, which looks like an agent that gives up.

Try it
  1. Deliberately cause three of these: name an agent "my agent", put two agents called critic in one team, and unset OPENAI_API_KEY before a run.
  2. Read each traceback and find the sentence that tells you the fix.
  3. If you have Ollama or any OpenAI-compatible server, point a client at it with a full model_info and run your hello world against it.
breaking things on purpose, while you are calm and know the cause, is how error messages become familiar. The third step also gives you a free model to iterate against.

Putting it all together

One script that uses everything above: a tool-using researcher, a reviewer, a termination pair, streaming output, token stats and saved state. Treat it as the template for your own first project.

research_team.py
import asyncio
import json

from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_core.model_context import BufferedChatCompletionContext
from autogen_ext.models.openai import OpenAIChatCompletionClient

# A stand-in for a real data source: a dict keyed by city.
CITY_DATA = {
    "cairo": {"population_m": 22.2, "timezone": "EET"},
    "riyadh": {"population_m": 7.7, "timezone": "AST"},
    "dubai": {"population_m": 3.7, "timezone": "GST"},
}

async def lookup_city(city: str) -> str:
    """Look up the population in millions and the timezone for a city."""
    record = CITY_DATA.get(city.strip().lower())
    if record is None:
        return f"No data for {city}. Known cities: {', '.join(sorted(CITY_DATA))}."
    return json.dumps(record)

async def main() -> None:
    model_client = OpenAIChatCompletionClient(model="gpt-4.1")

    researcher = AssistantAgent(
        "researcher",
        model_client=model_client,
        description="Looks up city facts with the lookup_city tool and drafts the answer.",
        system_message=(
            "You answer questions about cities. Use the lookup_city tool for every number "
            "you report, and never guess a figure. Address the reviewer's feedback when given."
        ),
        tools=[lookup_city],
        reflect_on_tool_use=True,
        max_tool_iterations=5,
        model_context=BufferedChatCompletionContext(buffer_size=12),
    )

    reviewer = AssistantAgent(
        "reviewer",
        model_client=model_client,
        description="Checks that every figure came from the tool and the answer is complete.",
        system_message=(
            "Check the draft. Every number must be traceable to a tool result, and the answer "
            "must cover every city asked about. Give one concrete correction if it does not. "
            "If the draft is correct and complete, reply with the single word APPROVE."
        ),
        model_context=BufferedChatCompletionContext(buffer_size=12),
    )

    termination = TextMentionTermination("APPROVE", sources=["reviewer"]) | MaxMessageTermination(12)
    team = RoundRobinGroupChat([researcher, reviewer], termination_condition=termination)

    result = await Console(
        team.run_stream(task="Compare Cairo, Riyadh and Dubai by population and timezone."),
        output_stats=True,
    )
    print("\nstop_reason:", result.stop_reason)
    print("usage:", model_client.total_usage())

    with open("team_state.json", "w") as f:
        json.dump(await team.save_state(), f)

    await model_client.close()

asyncio.run(main())

Walk through the decisions, because each one is a thing you now know rather than a thing copied from a sample.

The tool returns JSON and, when it fails, returns a useful sentence rather than raising. That is deliberate: a model reads the string, so an error message listing the known cities lets it recover on its own turn. max_tool_iterations=5 is there because three cities need three lookups, and the default of one would have stopped after the first. reflect_on_tool_use=True is there because the answer should be prose comparing three cities, not three raw JSON blobs.

The termination condition has the sources=["reviewer"] restriction so the researcher cannot end the run by quoting APPROVE in a draft — a small, real failure mode. MaxMessageTermination(12) is the bound, and it counts the task message. Both agents get a buffered context so a long disagreement cannot grow unboundedly, and both description fields are written properly, which becomes necessary the moment you switch to SelectorGroupChat.

Try it
  1. Run the script. Confirm the reviewer approves and read the token stats.
  2. Remove a city from CITY_DATA and run again. Watch the researcher handle the error string.
  3. Write the second script that loads team_state.json and asks "which of those three is furthest east?" without repeating the original task.
a multi-agent program with tools, review, bounded cost and durable state — which is the full beginner skill set. Step two is the one worth dwelling on: a tool that returns a helpful sentence on failure is the cheapest reliability improvement in agent engineering.

What you can now do, and what comes next

You can set up a pinned AutoGen environment and verify it without spending a token. You can create a model client for a hosted or a local model and describe an unrecognised model with model_info. You can build an agent with a system message and tools, and you know why max_tool_iterations and reflect_on_tool_use change its behaviour so much. You can put agents in a RoundRobinGroupChat, express "done" as a combined termination condition, and read stop_reason to find out what really happened. You can stream a run, count its tokens, bound its context, give it memory, save its state and resume it in another process. And you can read AutoGen's error messages and recognise a v0.2 or AG2 tutorial on sight.

What you have not touched is the rest of the framework. The next things to learn, roughly in order of how often they come up at work:

  • SelectorGroupChat and Swarm, for routing that is not a fixed order — including selector_func to skip the selection model call when the choice is deterministic, and handoffs= with HandoffTermination for explicit agent-to-agent transfer.
  • Workbenches and MCP, which is how an agent reaches real external systems. Remember "mcp<2" and the "only connect to trusted MCP servers" warning.
  • Code execution with CodeExecutorAgent and DockerCommandLineCodeExecutor, which became the default executor in 0.7.5 for good reason. Pass approval_func= or you will get a warning telling you code is running with no human oversight. The Docker guide covers the container side.
  • Structured output with output_content_type=MyPydanticModel, which turns an agent's reply into a validated object instead of prose you have to parse.
  • Agents as tools — AgentTool and TeamTool — remembering that they must not be called in parallel, so the parent's client needs parallel_tool_calls=False.
  • GraphFlow and DiGraphBuilder, for fan-out, fan-in and loops, still marked experimental.
  • The Core API: runtimes, AgentId, topics and subscriptions, and RoutedAgent with @message_handler. This is where multi-tenancy and event-driven designs live, and it is the senior guide's territory.
  • Observability and deployment: OpenTelemetry spans through an OTLP backend, one team per session behind FastAPI, state in a database.

For context on where AutoGen sits, read two or three neighbours. CrewAI takes a role-and-task view of the same problem. LangGraph makes the state machine explicit from the start, which is closer to GraphFlow than to RoundRobinGroupChat. Semantic Kernel is the other parent of the successor framework, so time spent there is directly useful. The OpenAI Agents SDK is the lighter-weight single-vendor comparison. And once anything you build runs for someone other than you, Langfuse is how you find out what it actually did.

A closing word on judgement. Learn AutoGen for what it teaches: the vocabulary of agents, teams, tools, handoffs and termination is now shared across the field, and the concepts transfer almost one-for-one to Microsoft Agent Framework. If you own AutoGen code at work, pin it carefully, keep up with security fixes yourself, and plan a migration with the official guide in hand rather than under time pressure. If you are starting something new, start it on the successor — and this guide taught you most of what you need for that too.

Next in this series: the Mid-level guide, which picks up at SelectorGroupChat, workbenches, structured output, state in a real service and debugging beyond Console.

Sources