This is part one of three. It covers everything you need to build working multi-agent programs with Microsoft AutoGen, starting from a machine that has never had it installed. By the end you can create an agent backed by a language model, give it Python functions to call, put two or three agents in a team that stops when it should, stream the conversation to your terminal, save and restore a run, and read the error messages AutoGen produces when something is wrong.
One thing has to be said before the first command, because it changes how you should read the rest: AutoGen is in maintenance mode. The README on github.com/microsoft/autogen states that it "will not receive new features or enhancements and is community managed going forward", and that "new users should start with Microsoft Agent Framework". AutoGen 0.7.5 is still a working, widely deployed framework, and the ideas in it — agents, teams, termination conditions, tool loops, handoffs — are the ideas the successor uses too. So learning it is not wasted. But you should know from the first page that this is a tool you learn to maintain and migrate, not one you bet a brand-new three-year project on. The second section covers that in detail, and every later section points out where the successor differs.
Each section ends with a Try it task. Do them as you go. Multi-agent code is deceptively easy to read and surprisingly easy to get wrong, and the only reliable cure is to watch your own run stop for the wrong reason once.
What AutoGen is, and the problem it solves
AutoGen is a Python framework for building applications in which one or more agents — programs that wrap a language model, a set of tools and a bit of memory — work on a task, and where some of them talk to each other to get it done. You describe the agents and how they take turns; AutoGen runs the loop, routes the messages, calls the tools, and tells you why it stopped.
To see what it buys you, write the thing it replaces. Suppose you want a model to answer questions about a city's weather. Without a framework you call the chat completions endpoint, read the response, notice that the model asked to call a function, find that function in a dictionary of your own making, call it, append the result to the message list in exactly the shape the provider expects, call the endpoint again, and check whether the model is done. That is forty or fifty lines of plumbing, none of it your idea. Now add a second model that reviews the first one's answer: you also own a turn-taking loop, a rule for when the conversation is finished, a cap so a disagreement cannot run forever, and a way to see what happened when it goes wrong at three in the morning.
Every line in that paragraph is a thing AutoGen owns for you. The tool-calling round trip becomes tools=[get_weather]. The turn-taking loop becomes RoundRobinGroupChat([writer, critic]). The rule for finishing becomes TextMentionTermination("APPROVE") | MaxMessageTermination(10). The visibility becomes await Console(team.run_stream(task=...)).
That diagram is the whole shape of an AgentChat program, and almost everything in this guide is an elaboration of one of those five boxes.
It is worth knowing what came before, because you will meet the ancestor in search results. AutoGen started in 2023 as a research project whose central object was ConversableAgent and whose central verb was initiate_chat. That API — usually called v0.2 — was quick to demo and hard to reason about: agents replied through registered reply functions, the control flow lived inside the library, and configuration arrived as an llm_config dictionary. Microsoft rewrote it in early 2025. The v0.4 line, which 0.7.5 belongs to, split the framework into an event-driven runtime and a high-level task API, made every agent explicitly stateful, made asynchrony first class, and made the stopping rule a real object you can combine with | and &.
- Write out, in plain language, a task you would like two language models to do together — one producing something, one checking it.
- Underneath, write the rule that tells you the task is finished. Be exact: a word one of them says, a number of turns, a time limit.
- Keep the note. It becomes your termination condition later in this guide.
Status first: maintenance mode, and what it means for your code
Treat this as part of the installation, not as a footnote. A framework's support status changes which habits are correct.
The facts, as of this guide's research date. The maintenance-mode banner went on the repository README on 2 October 2025 and was refreshed in April 2026. The last release published to PyPI is 0.7.5, dated 30 September 2025; commits after it are bug fixes, security fixes and documentation. The repository is not archived, so issues and pull requests still move, but no new capability is coming. The named successor is Microsoft Agent Framework, built by the AutoGen and Semantic Kernel teams together, which is at 1.x (the Python agent-framework package was at 1.19.0 in September 2026). Microsoft publishes an official AutoGen-to-Agent-Framework migration guide on Microsoft Learn.
Three practical consequences follow, and they shape the advice in every later section.
Pin everything, exactly. A maintained library absorbs the churn of its dependencies for you. A library in maintenance mode does not. The clearest live example is the Model Context Protocol extra: autogen-ext[mcp]==0.7.5 declares mcp>=1.11.0 with no upper bound, so a fresh install today pulls mcp 2.x, and the first import fails with ImportError: cannot import name 'RequestContext' from 'mcp.shared.context'. Nothing is wrong with your code; a dependency moved and nobody updated the bound. You fix it by pinning mcp<2 yourself. Assume more of these over time, and use a lock file from your very first project.
Read the release notes before any upgrade, even a minor one. AutoGen's 0.x minors carried real breaking changes — 0.4 to 0.5 to 0.6 to 0.7 each moved something. Pin autogen-agentchat, autogen-core and autogen-ext to the same exact version, which is what their own metadata expects: autogen-agentchat pins autogen-core to an exact match.
Learn the concepts with an eye on the mapping. The migration guide gives a direct table, and it is short enough to internalise now. AssistantAgent becomes Agent. OpenAIChatCompletionClient keeps its name. FunctionTool becomes a @tool decorator. AgentTool(agent) becomes agent.as_tool(). RoundRobinGroupChat becomes SequentialBuilder. MagenticOneGroupChat becomes MagenticBuilder. GraphFlow becomes WorkflowBuilder. Model context and state become AgentSession. One behavioural difference is worth memorising because it will bite you in both directions: AutoGen's AssistantAgent is effectively single-turn unless you raise max_tool_iterations, whereas the successor's Agent is multi-turn by default.
- Open the repository README at
github.com/microsoft/autogenand find the maintenance notice yourself. - Open the releases page and read the notes for the most recent release.
- Write down one thing in those notes you do not understand yet, and look for it later in this guide.
Three packages and one very confusing name
AutoGen ships as three Python packages, and understanding the split tells you where to look for anything.
| Package | Import name | What lives in it |
|---|---|---|
autogen-core |
autogen_core |
The event-driven runtime, the base types, message passing, agent identity, tool and memory interfaces |
autogen-agentchat |
autogen_agentchat |
The high-level, task-driven API: preset agents, teams, termination conditions, the Console UI |
autogen-ext |
autogen_ext |
Everything that touches the outside world: model clients, code executors, memory backends, MCP, extra runtimes |
Beginners should live almost entirely in AgentChat, reaching into ext for a model client and into core for a couple of helper types. Core's own API — runtimes, topics, subscriptions, routed agents — is a different and lower-level way to write the same kinds of programs, and it belongs at mid and senior level. The three-layer split exists so that AgentChat can stay opinionated while Core stays general.
Alongside the libraries are three developer tools you will see mentioned, none of them needed to learn the framework. AutoGen Studio is a low-code browser UI for assembling teams by clicking; its own documentation says it is "not meant to be a production-ready app", and the stable release 0.4.2.2 requires the autogen-* packages below 0.6, so it must live in a separate virtual environment from your 0.7.x code. Magentic-One is a prebuilt generalist team plus an m1 command line, installed from magentic-one-cli. AgBench is a benchmarking harness.
Now the name problem, which costs beginners more time than any technical issue in this guide.
Microsoft AutoGen — what this guide teaches
- Install:
autogen-agentchat,autogen-core,autogen-ext - Imports use underscores:
from autogen_agentchat.agents import AssistantAgent - Docs:
microsoft.github.io/autogen/stable/ - Objects:
AssistantAgent,RoundRobinGroupChat,await agent.run(...)
Not this guide
- PyPI
autogenandag2belong to AG2, the community fork at ag2ai/ag2 import autogenis always v0.2-style or AG2 code- Docs:
docs.ag2.ai— a different, v0.2-descended API - Objects:
ConversableAgent,initiate_chat,llm_config
There is a third name in the mess. pyautogen was the original distribution, and the migration guide states that Microsoft "no longer [has] admin access to the pyautogen PyPI package, and the releases from that package are no longer from Microsoft since version 0.2.34". Today pyautogen 0.10.0 is a thin proxy depending on autogen-agentchat, so installing it is harmless but pointless. The honest rule is the import rule: if the code says import autogen, it is not the framework you installed.
ConversableAgent, initiate_chat, register_reply, llm_config, config_list or cache_seed. A hit on any of them means the page predates the rewrite or belongs to the fork, and you should close it. A page that says AssistantAgent, model_client= and await is written for what you have.
- Search the web for "autogen tutorial" and open the first three results.
- Apply the triage above to each one and label it current, v0.2 or AG2.
- Note how many of the three you would have followed blindly.
AttributeError on ConversableAgent.
The mental model: four nouns
Everything in AgentChat is built from four nouns. Learn them in this order, because each one depends on the ones before it.
The model client is the thing that talks to a language model. Its interface is ChatCompletionClient, defined in autogen_core.models, and the concrete implementations live in autogen_ext.models.* — one for OpenAI and Azure OpenAI, one for Anthropic, one for Ollama, one for llama.cpp, one for Azure AI, an adapter for Semantic Kernel, and a ReplayChatCompletionClient that returns canned strings and is perfect for tests. A client exposes create() and create_stream() for calls, count_tokens() and remaining_tokens() for budgeting, actual_usage() and total_usage() for cost, a model_info property describing what the model can do, and close(). You create one client and share it between agents.
The agent has a name that must be a valid Python identifier, a description that teams use to decide who should speak, and the behaviour you configure: a system message, tools, memory, a view of the conversation history. The important property, and the one beginners most often get wrong, is that agents are stateful. An agent remembers the conversation across calls. So when you call it again you pass only the new message, never the whole transcript — passing the transcript duplicates everything in its context and doubles your token bill for no benefit.
The team is a group of agents sharing one message thread, driven by a group-chat manager. Four presets matter at this level, and RoundRobinGroupChat is the only one you need today. It gives agents a fixed turn order. SelectorGroupChat asks a model to pick the next speaker each turn. Swarm lets agents hand off to each other explicitly. MagenticOneGroupChat wraps an orchestrator that plans, tracks whether progress has stalled, and re-plans. A fifth, GraphFlow, lets you draw the control flow as a directed graph and is still documented as experimental.
The termination condition decides when a run stops. It is a stateful callable, checked after each agent response against only the new messages, which returns a StopMessage or None. You combine conditions with | for "or" and & for "and", and each one resets itself after a run finishes.
Two return types carry results. Response, returned by an agent's low-level on_messages, has chat_message (the final message) and inner_messages (the events that led to it). TaskResult, returned by run() and yielded last by run_stream(), has messages and — the field you will read most often in your life — stop_reason, a human-readable string saying exactly why the loop ended.
Messages come in two families, and knowing which is which stops a lot of confusion when you start streaming. Chat messages travel between agents: TextMessage, MultiModalMessage, StopMessage, HandoffMessage, ToolCallSummaryMessage and the generic StructuredMessage[T]. Events are internal to one agent and exist so you can watch it think: ToolCallRequestEvent, ToolCallExecutionEvent, MemoryQueryEvent, UserInputRequestedEvent, ModelClientStreamingChunkEvent, ThoughtEvent and SelectSpeakerEvent. Every message carries source, models_usage, metadata, id and created_at, and offers to_text().
- Say the four nouns out loud, each with one sentence of your own words.
- For your task from the first section, name the agents, say which team preset fits, and state the termination condition.
- Decide which model you will use, and whether it supports function calling.
tools= on a model whose model_info says function_calling: False raises ValueError: The model does not support function calling. before it ever reaches the network.
Installing AutoGen and checking it works
Use a virtual environment. AutoGen has a wide dependency surface and version pinning matters more here than in a maintained library, so a per-project environment is not optional advice.
python3 -m venv .venv
source .venv/bin/activate # Linux and macOS
# Windows cmd: .venv\Scripts\activate.bat
# Windows PowerShell: .venv\Scripts\Activate.ps1
If you prefer conda, conda create -n autogen python=3.12 && conda activate autogen is equivalent. Python 3.10 is the minimum the packages declare; 3.12 is a safe choice.
Then install. The first line is the official README's; the second is what you should commit.
pip install -U "autogen-agentchat" "autogen-ext[openai]"
# For reproducibility, pin the exact versions instead:
pip install "autogen-agentchat==0.7.5" "autogen-ext[openai]==0.7.5"
autogen-core arrives automatically, because autogen-agentchat pins it to the same exact version. The square brackets are pip extras: autogen-ext is a bag of optional integrations and you install only the ones you need. The ones worth knowing early are openai, anthropic, ollama, gemini, azure and llama-cpp for models; mcp, docker and jupyter-executor for tools and code execution; redis, chromadb, mem0 and diskcache for memory and caching; magentic-one, web-surfer and file-surfer for the prebuilt agents; and grpc and rich for the distributed runtime and prettier terminal output. In PowerShell, keep the quotes around the whole argument or the shell eats the brackets.
Set your key as an environment variable. OpenAIChatCompletionClient reads OPENAI_API_KEY by itself, and the Anthropic client reads ANTHROPIC_API_KEY.
export OPENAI_API_KEY=sk-... # PowerShell: $env:OPENAI_API_KEY="sk-..."
Now verify, in two steps. First that the packages are importable and the versions agree:
python -c "import autogen_agentchat, autogen_core; print(autogen_agentchat.__version__, autogen_core.__version__)"
# -> 0.7.5 0.7.5
pip show autogen-agentchat autogen-core autogen-ext
Second — and do this before you spend a single token — run the framework end to end with no model at all. ReplayChatCompletionClient is a real client that returns the strings you give it in order. It proves your installation, your async setup and your mental model are all correct, and it costs nothing.
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.replay import ReplayChatCompletionClient
async def main():
agent = AssistantAgent("assistant", model_client=ReplayChatCompletionClient(["Hello World!"]))
result = await agent.run(task="Say hello")
print(result.messages[-1].content) # Hello World!
asyncio.run(main())
Three per-platform notes save a lot of grief. On macOS and Linux use python3; Docker (Docker Desktop or Colima on macOS, Docker Engine on Linux) is needed for the Docker code executor, and autogen-ext[web-surfer] also needs playwright install. On Windows, the local code executor needs the Proactor event loop, and the source warns "The current event loop policy is not WindowsProactorEventLoopPolicy…"; fix it with asyncio.set_event_loop_policy(asyncio.WindowsProactorEventLoopPolicy()). In Jupyter or IPython do not call asyncio.run(...) — the notebook already runs an event loop, and you will get RuntimeError: asyncio.run() cannot be called from a running event loop. Use top-level await.
- Create the virtual environment, install the pinned versions, and run the two verification commands.
- Save and run
smoke_test.py. Confirm it printsHello World!with no API key set. - Run
pip freeze > requirements.txtand open the file.
Your first real agent
Now the same program against a real model. This is the official hello world, and it is worth reading line by line because every later example is a variation of it.
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_ext.models.openai import OpenAIChatCompletionClient
async def main() -> None:
model_client = OpenAIChatCompletionClient(model="gpt-4.1")
agent = AssistantAgent("assistant", model_client=model_client)
print(await agent.run(task="Say 'Hello World!'"))
await model_client.close()
asyncio.run(main())
Four things are happening. The client is created with a model name and no key, because it falls back to OPENAI_API_KEY. The agent is created with a name — "assistant" is a valid Python identifier, which is required — and the client. run(task=...) sends one message and returns a TaskResult, which is why printing it gives you the whole structure rather than just the text. And close() releases the client's HTTP connections; forget it and Python will complain about unclosed sessions at exit.
AssistantAgent has more knobs than any other class you will touch, so here are the ones that matter at this level, with their real defaults.
| Parameter | Default | What it does |
|---|---|---|
system_message |
A built-in instruction ending "Reply with TERMINATE when the task has been completed." | The persona and the rules. Change it for almost every real agent |
description |
"An agent that provides assistance with ability to use tools." | How a team decides this agent should speak. Worth writing properly once you have more than one agent |
tools |
None |
Python callables the model may invoke |
max_tool_iterations |
1 |
How many model-then-tool rounds happen inside one turn |
reflect_on_tool_use |
None |
Whether the model gets a chance to summarise the tool result in prose |
model_client_stream |
False |
Whether token chunks are emitted as events |
output_content_type |
None |
A Pydantic model for structured output |
memory |
None |
Memory stores queried before each model call |
model_context |
None (unbounded) |
Which slice of history is sent to the model |
Note that default system message. The stock AssistantAgent has been told to say TERMINATE when it is finished, which is why so many examples pair it with TextMentionTermination("TERMINATE"). If you replace the system message — and you should, for any agent with a real job — you also take responsibility for whatever word your termination condition is looking for. A termination condition watching for a word your agent was never told to say produces a run that goes to its message cap every time, and the only clue is stop_reason.
There are three ways to run things, and they differ only in how much you see:
result = await agent.run(task="...")gives you theTaskResultat the end and nothing in between.async for message in agent.run_stream(task="..."): ...yields each message and event as it happens, with theTaskResultlast.await Console(agent.run_stream(task="..."))prints that stream to your terminal in a readable form.Consolecomes fromautogen_agentchat.ui, andConsole(..., output_stats=True)adds token counts.
task can be a plain string, a BaseChatMessage, or a list of them. Both run and run_stream also take cancellation_token= to abort and output_task_messages= (default True) to control whether your own task is echoed back in the result.
- Run
hello.pyand read the printedTaskResult. Findstop_reasonin it. - Replace the
printwithawait Console(agent.run_stream(task="Explain what a container is in three sentences."))and importConsolefromautogen_agentchat.ui. - Give the agent
system_message="You answer only in bullet points, in at most 40 words."and run it again.
Console in your scratch scripts permanently; debugging an agent you cannot see is guesswork.
Giving an agent tools
A tool is how an agent does something other than produce text. In AutoGen, a tool is any Python function — synchronous or asynchronous — with type hints on its parameters and a docstring. You pass the function itself and AutoGen wraps it in a FunctionTool for you. The type hints become the schema the model sees, and the docstring becomes the description the model uses to decide whether to call it, so both are load-bearing rather than decorative.
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient
async def get_weather(city: str) -> str:
"""Get the weather for a given city."""
return f"The weather in {city} is 73 degrees and Sunny."
async def main() -> None:
model_client = OpenAIChatCompletionClient(model="gpt-4.1")
agent = AssistantAgent(
"weather_agent",
model_client=model_client,
tools=[get_weather],
reflect_on_tool_use=True,
)
await Console(agent.run_stream(task="What is the weather in Dubai?"))
await model_client.close()
asyncio.run(main())
Run it with Console and read the stream carefully, because it shows the tool protocol in full: a ToolCallRequestEvent where the model asks for get_weather(city="Dubai"), a ToolCallExecutionEvent carrying what your function returned, and then a final message.
What that final message is depends on reflect_on_tool_use. With it set to True, the tool result goes back to the model and the agent's final message is the model's prose. With it left off, the agent's final message is a ToolCallSummaryMessage containing the raw tool result formatted with tool_call_summary_format, which defaults to "{result}". Neither is wrong. Reflection reads better and costs an extra model call; the summary is cheaper, faster and exactly what you want when the tool's output is already the answer. The default is subtle: reflect_on_tool_use=None resolves to True when output_content_type is set and False otherwise.
AssistantAgent.max_tool_iterations defaults to 1, which means a single model-then-tool round per turn. If your task needs the agent to call a tool, look at the result, and then call another tool, it will not do it — it will stop after the first. Raise the limit explicitly, for example max_tool_iterations=10. This parameter arrived in 0.6.2, so a great many tutorials written before it do not mention it, and their authors worked around the limit by wrapping the agent in a team.
Two rules about tools are worth learning now rather than discovering through an exception. First, an agent takes tools= or workbench=, never both; passing both raises ValueError: Tools cannot be used with a workbench. A workbench is a set of tools that share state and resources — StaticWorkbench for a fixed list, McpWorkbench for a Model Context Protocol server — and it is the mid-level topic. Second, the model must support function calling. If model_info says function_calling: False, passing tools= raises ValueError: The model does not support function calling. straight away.
Tools are also where security enters. A tool runs with your program's privileges, and its arguments come from a model that may be acting on text a stranger wrote. Validate them inside the function as you would an HTTP request body: check ranges, restrict paths, never interpolate a model-supplied string into a shell command or SQL statement. The framework does not do this for you, and the same caution applies with more force to MCP servers — the README's warning is "Only connect to trusted MCP servers." See the MCP guide.
- Run
weather.pyand identify the two tool events in theConsoleoutput. - Remove
reflect_on_tool_use=Trueand run it again. Compare the final messages. - Add a second tool,
def convert_f_to_c(f: float) -> float, and ask for the weather in Dubai in Celsius. Then setmax_tool_iterations=5and ask again.
Two agents in a team, and knowing when to stop
A team is a group of agents sharing one message thread. The simplest and most predictable is RoundRobinGroupChat, which gives them a fixed turn order. The classic shape is a producer and a reviewer.
import asyncio
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_ext.models.openai import OpenAIChatCompletionClient
async def main() -> None:
model_client = OpenAIChatCompletionClient(model="gpt-4.1")
primary = AssistantAgent(
"primary",
model_client=model_client,
description="Writes the first draft.",
system_message="You are a helpful writer. Write short, concrete prose.",
)
critic = AssistantAgent(
"critic",
model_client=model_client,
description="Reviews drafts and approves them.",
system_message=(
"Review the draft and give specific, actionable feedback. "
"When the draft needs no further change, reply with the single word APPROVE."
),
)
termination = TextMentionTermination("APPROVE") | MaxMessageTermination(10)
team = RoundRobinGroupChat([primary, critic], termination_condition=termination)
result = await Console(team.run_stream(task="Write a short poem about the Red Sea in autumn."))
print(result.stop_reason)
await team.reset()
await model_client.close()
asyncio.run(main())
Read the termination condition as English: stop when someone says APPROVE, or when ten messages have gone by. The second half is not optional pessimism, it is the safety net. A team with no termination condition and no max_turns runs until your budget or the provider stops it, and two polite agents can disagree about a poem for a remarkably long time. Make the pairing a habit: one condition that expresses what you actually want, one that bounds the damage.
Here are the built-in conditions, all from autogen_agentchat.conditions:
| Condition | Stops when |
|---|---|
MaxMessageTermination(max_messages, include_agent_event=False) |
A message count is reached |
TextMentionTermination(text, sources=None) |
Some text appears in a message, optionally only from named sources |
TextMessageTermination(source=None) |
A TextMessage is produced, optionally by a named source |
StopMessageTermination() |
An agent produces a StopMessage |
TokenUsageTermination(max_total_token=..., max_prompt_token=..., max_completion_token=...) |
A token budget is spent |
TimeoutTermination(timeout_seconds) |
Wall-clock time runs out |
HandoffTermination(target) |
A handoff to a given target happens, typically "user" |
SourceMatchTermination(sources) |
One of the named agents has spoken |
FunctionCallTermination(function_name) |
A particular tool is called |
ExternalTermination() |
You call .set() on it from elsewhere |
FunctionalTermination(func) |
Your own function says so |
One counting detail trips up nearly everybody: MaxMessageTermination counts the task message. With MaxMessageTermination(3) you get your task plus two agent replies, and the stop reason reads "Maximum number of messages 3 reached, current message count: 3". If you want N replies, ask for N+1.
The other behaviour to internalise is run and resume. Calling team.run() again — with no task, or with a new one — continues from where the previous run stopped; the agents still remember everything. await team.reset() is what clears the state and gives you a clean slate. This is a feature, and it is also the reason a long-running script slowly gets more expensive if you never reset: the context keeps growing.
You will meet the other presets at mid level, but a one-line map helps now. SelectorGroupChat spends an extra model call each turn asking which agent should speak next, needs at least two participants, and by default will not let the same agent speak twice in a row. Swarm routes by explicit handoffs, where the next speaker is whoever the most recent HandoffMessage named. MagenticOneGroupChat runs an orchestrator with a task and progress ledger, defaults to twenty turns, and gives up after three stalls. Since 0.7.1 a team can itself be a participant of the first two, which is how you nest a sub-team.
- Run
team.pyand read the printedstop_reason. - Delete
| MaxMessageTermination(10), changeAPPROVEin the critic's system message toPERFECT, but leave the condition watching forAPPROVE. Run it with a lowmax_turns=4on the team so you can stop it safely. - Restore the pair, then call
team.run()a second time withtask="Now make it two lines shorter."without resetting.
Seeing what happened: streaming, events and state
Debugging agents is mostly a visibility problem, so spend ten minutes getting good at this and you will save hours later.
run_stream yields every chat message and every event in order, with the TaskResult last. Wrapping it in Console is the quick option; iterating it yourself is how you build your own UI or log.
from autogen_agentchat.base import TaskResult
async for item in team.run_stream(task="Write a haiku about Cairo traffic."):
if isinstance(item, TaskResult):
print("STOPPED:", item.stop_reason)
else:
print(f"[{item.source}] {type(item).__name__}: {item.to_text()[:120]}")
Setting model_client_stream=True on an AssistantAgent adds ModelClientStreamingChunkEvents, so you see tokens as they arrive rather than complete messages. That is what you want behind a chat interface, and it is noise in a log.
Token usage is attached to the messages themselves. Every message carries models_usage, a RequestUsage with prompt_tokens and completion_tokens, and the client keeps running totals you can read with total_usage() and actual_usage(). Console(..., output_stats=True) prints them for you. Look at these numbers on your very first team run; the shape of the bill is rarely what people guess, because each agent re-reads the whole shared thread on its turn.
State is the other half of visibility. save_state() on a team or an agent returns a JSON-serialisable dictionary, and load_state() restores it.
import json
state = await team.save_state()
with open("state.json", "w") as f:
json.dump(state, f)
# later, in another process
with open("state.json") as f:
await new_team.load_state(json.load(f))
For an AssistantAgent, that state is essentially its model context — the conversation it remembers. This is how you build a web application on top of AgentChat: one team per user session, state written to Postgres or Redis between requests, loaded again on the next one. You cannot load_state while a team is running, and trying raises RuntimeError: The team cannot be loaded while it is running.
Two related constraints surprise people building their first service. A team instance cannot run twice concurrently: a second overlapping run() raises ValueError: The team is already running, it cannot run again until it is stopped. Never share one team object across HTTP requests. And for controlled stopping, prefer ExternalTermination().set(), which lets the current agent finish its turn, over a CancellationToken, which aborts immediately and can leave state inconsistent.
Finally, plain Python logging works: autogen_core exposes TRACE_LOGGER_NAME for human-readable debug output and EVENT_LOGGER_NAME for structured events. The runtime is also instrumented with OpenTelemetry, emitting GenAI spans named create_agent, invoke_agent and execute_tool, so any OTLP backend can show you a trace of a run — that is how you would wire it to Langfuse. AUTOGEN_DISABLE_RUNTIME_TRACING=true switches the runtime spans off.
- Rewrite your team run with the explicit
async forloop above and watch the event types scroll past. - Add
output_stats=Trueto aConsolecall and note the total tokens for a two-agent, six-message run. - Save the team state to
state.json, open the file, then load it into a freshly constructed team and continue the conversation.
state.json by hand is the moment the framework stops being magic: it is a conversation, a few counters, and your configuration. Anything you can save you can inspect, diff and test.
Putting a human in the loop
Sometimes the right next speaker is a person. AutoGen gives you two ways to arrange that, and choosing wrongly is a classic first-project mistake.
The direct way is UserProxyAgent, an agent whose turn is taken by a human. Its signature is UserProxyAgent(name, *, description="A human user", input_func=None) and input_func defaults to Python's input().
from autogen_agentchat.agents import AssistantAgent, UserProxyAgent
from autogen_agentchat.conditions import TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
assistant = AssistantAgent("assistant", model_client=model_client)
user = UserProxyAgent("user_proxy")
team = RoundRobinGroupChat(
[assistant, user],
termination_condition=TextMentionTermination("APPROVE", sources=["user_proxy"]),
)
await Console(team.run_stream(task="Draft a two-sentence summary of our data residency policy."))
This is perfect for a terminal script and wrong for a web application, for a reason worth understanding. While UserProxyAgent waits for input it blocks the running team, and a running team cannot be saved. So a web request that is waiting for a human is a request holding a live team object with unsaveable state, which is exactly the thing you cannot do at scale.
The pattern that does work is to make "we need the human" a reason to stop. End the run, return to your application, and call run() again with the person's reply when it arrives. HandoffTermination(target="user") and a plain TextMentionTermination both do this, and because resuming a team is normal, the second run() picks up precisely where the first stopped. You can watch for UserInputRequestedEvent in the stream to know a request for input is coming.
UserProxyAgent. Asynchronous human — someone who will reply within the next hour, through a web form or a chat app — terminate, persist the state, and resume. The second shape also survives a process restart, which the first does not.
- Run
human.pyand reply a couple of times, finishing withAPPROVE. - While it is waiting for your input, try to picture what
await team.save_state()would have to capture. Then read the paragraph above again. - Rewrite it without
UserProxyAgent: terminate on a word, print the draft, useinput()in your own code, and callteam.run(task=your_reply).
Context, memory and keeping the bill down
By default an agent sends its entire conversation history to the model on every turn. That is the UnboundedChatCompletionContext, and for a short task it is exactly right. For a long one it is a slow-motion cost problem: every turn is longer than the last, and eventually you hit the model's context limit and the run fails.
A model context is the configurable answer to "which part of history does the model actually see". There are four, all in autogen_core.model_context:
| Model context | Sends |
|---|---|
UnboundedChatCompletionContext |
Everything. The default |
BufferedChatCompletionContext(buffer_size=N) |
The last N messages |
HeadAndTailChatCompletionContext(head_size=, tail_size=) |
The first few and the last few, dropping the middle |
TokenLimitedChatCompletionContext(model_client, token_limit=) |
As much as fits in a token budget |
HeadAndTailChatCompletionContext encodes a real insight: the beginning of a conversation holds the task and the constraints, the end holds the current state, and the middle is the most droppable part.
Memory is the complementary idea. A memory store is queried before each model call and its results are injected into the context, which lets an agent carry facts across conversations rather than within one. The interface is autogen_core.memory.Memory, and the simplest implementation is ListMemory.
from autogen_core.memory import ListMemory, MemoryContent, MemoryMimeType
from autogen_core.model_context import BufferedChatCompletionContext
mem = ListMemory()
await mem.add(MemoryContent(content="User prefers metric units", mime_type=MemoryMimeType.TEXT))
agent = AssistantAgent(
"assistant",
model_client=model_client,
model_context=BufferedChatCompletionContext(buffer_size=10),
memory=[mem],
)
When memory fires you will see a MemoryQueryEvent in the stream, which is how you confirm it is doing anything. Beyond ListMemory, autogen_ext offers vector and service-backed stores — ChromaDBVectorMemory, RedisMemory, Mem0Memory and a canvas memory — each behind its own pip extra.
While we are on cost, collect the levers in one place, because this is the question your manager will ask. Cap the conversation with MaxMessageTermination or max_turns. Cap the spend directly with TokenUsageTermination. Cap the context with a buffered or token-limited model context. Avoid SelectorGroupChat when routing is deterministic, since it pays for an extra model call every turn just to choose a speaker. Leave reflect_on_tool_use off when the tool's raw output is already the answer, since reflection is another model call. And cache repeated calls: model-response caching has been off by default since v0.4, and you turn it on by wrapping your client.
from autogen_ext.models.cache import ChatCompletionCache, CHAT_CACHE_VALUE_TYPE
from autogen_ext.cache_store.diskcache import DiskCacheStore
from diskcache import Cache
cached_client = ChatCompletionCache(
model_client,
DiskCacheStore[CHAT_CACHE_VALUE_TYPE](Cache("/tmp/autogen-cache")),
)
That needs pip install "autogen-ext[diskcache]", and there is a Redis equivalent in autogen_ext.cache_store.redis. While you are developing and running the same prompt twenty times, a cache is the single biggest saving available to you.
- Run a six-turn team conversation with
Console(..., output_stats=True)and record the total tokens. - Give both agents
model_context=BufferedChatCompletionContext(buffer_size=4)and run the same task again. Compare. - Add a
ListMemoryentry stating a preference, ask a question that should respect it, and find theMemoryQueryEventin the stream.
Configuration, and the errors you will actually see
Two configuration topics matter at this level: telling AutoGen about a model it does not recognise, and saving a setup as JSON.
The first is model_info, a small dictionary describing what a model can do. The OpenAI client knows OpenAI's own models, so for gpt-4.1 you pass nothing. The moment you point it at something else — a local model through Ollama's OpenAI-compatible endpoint, a vLLM or LM Studio server, a LiteLLM proxy, Gemini's OpenAI-compatible endpoint — it has no idea, and you must say.
from autogen_ext.models.openai import OpenAIChatCompletionClient
client = OpenAIChatCompletionClient(
model="llama3.2",
base_url="http://localhost:11434/v1",
api_key="placeholder",
model_info={
"vision": False,
"function_calling": True,
"json_output": False,
"family": "unknown",
"structured_output": True,
},
)
Omit it and you get ValueError: model_info is required when model name is not a valid OpenAI model. Omit just the structured_output key and you get UserWarning: Missing required field 'structured_output' in ModelInfo. This field will be required in a future version of AutoGen. — a warning today, an error later, so fill it in. Lie about function_calling and your tool calls fail at the provider instead of in AutoGen, which is a much worse place to debug. family takes a value from ModelFamily or the string "unknown". Note also that model_capabilities= is the deprecated predecessor of model_info and the two are mutually exclusive; if you see it in a tutorial, rename it.
A direct Ollama client also exists at autogen_ext.models.ollama.OllamaChatCompletionClient, taking a host key, which is tidier than the OpenAI-compatible endpoint. The Ollama guide covers the server side.
The second topic is component configuration. Every AutoGen component can serialise itself to JSON with dump_component() and be rebuilt with load_component().
cfg = team.dump_component()
json_str = cfg.model_dump_json()
team2 = RoundRobinGroupChat.load_component(cfg)
This is how AutoGen Studio stores teams, and it keeps configuration out of code. Two warnings come with it. Callables are not serialisable — a selector_func, a GraphFlow lambda condition and a FunctionTool cannot survive the round trip. And dump_component() writes an explicitly passed api_key into the JSON, so rely on environment variables and never commit a dumped config. There is a sharper point too: in released 0.7.5, load_component() imports whatever module path the provider string names, so loading component JSON from an untrusted source is remote code execution. A namespace restriction exists on main but is in no PyPI release. Treat component JSON as code.
Now the errors. AutoGen's messages are unusually good — most of them name the fix — so learning to read them is quick.
| Message | What it means |
|---|---|
openai.OpenAIError: Missing credentials. Please pass an api_key… |
No key. export OPENAI_API_KEY=sk-... or pass api_key= |
openai.AuthenticationError: Error code: 401 … 'code': 'invalid_api_key' |
The key is wrong or revoked |
openai.APIConnectionError: Connection error. |
Wrong base_url, server down, or TLS interception by a corporate proxy. Check the URL and your CA bundle |
ValueError: The agent name must be a valid Python identifier. |
You wrote "my agent" or "my-agent". Use my_agent |
ValueError: The participant names must be unique. |
Two agents in one team share a name |
ValueError: The model does not support function calling. |
tools= on a model whose model_info says otherwise |
ValueError: Tools cannot be used with a workbench. |
You passed both. Pick one |
ValueError: The team is already running, it cannot run again until it is stopped. |
Concurrent run() on one team object. One team per session |
RuntimeError: The team cannot be loaded while it is running. |
load_state mid-run |
ValueError: At least two participants are required for SelectorGroupChat. |
Use RoundRobinGroupChat for one agent |
RuntimeError: asyncio.run() cannot be called from a running event loop |
asyncio.run inside Jupyter. Use top-level await |
ModuleNotFoundError: No module named 'anthropic' (or ollama, chromadb, playwright) |
A missing pip extra. Install "autogen-ext[anthropic]" and so on |
ImportError: cannot import name 'RequestContext' from 'mcp.shared.context' |
mcp 2.x with autogen-ext 0.7.5. Pin "mcp<2" |
ModuleNotFoundError: No module named 'autogen' or AttributeError on ConversableAgent |
You are following a v0.2 or AG2 tutorial |
Four non-error pitfalls round out the list, and each one produces behaviour rather than a traceback, which makes them harder. A team with no termination condition and no max_turns runs until something external stops it. Passing the full history into agent.run() each time duplicates context, because agents are already stateful. Forgetting await model_client.close() leaves HTTP connections open and produces warnings at exit. And max_tool_iterations=1 means one tool round, which looks like an agent that gives up.
- Deliberately cause three of these: name an agent
"my agent", put two agents calledcriticin one team, and unsetOPENAI_API_KEYbefore a run. - Read each traceback and find the sentence that tells you the fix.
- If you have Ollama or any OpenAI-compatible server, point a client at it with a full
model_infoand run your hello world against it.
Putting it all together
One script that uses everything above: a tool-using researcher, a reviewer, a termination pair, streaming output, token stats and saved state. Treat it as the template for your own first project.
import asyncio
import json
from autogen_agentchat.agents import AssistantAgent
from autogen_agentchat.conditions import MaxMessageTermination, TextMentionTermination
from autogen_agentchat.teams import RoundRobinGroupChat
from autogen_agentchat.ui import Console
from autogen_core.model_context import BufferedChatCompletionContext
from autogen_ext.models.openai import OpenAIChatCompletionClient
# A stand-in for a real data source: a dict keyed by city.
CITY_DATA = {
"cairo": {"population_m": 22.2, "timezone": "EET"},
"riyadh": {"population_m": 7.7, "timezone": "AST"},
"dubai": {"population_m": 3.7, "timezone": "GST"},
}
async def lookup_city(city: str) -> str:
"""Look up the population in millions and the timezone for a city."""
record = CITY_DATA.get(city.strip().lower())
if record is None:
return f"No data for {city}. Known cities: {', '.join(sorted(CITY_DATA))}."
return json.dumps(record)
async def main() -> None:
model_client = OpenAIChatCompletionClient(model="gpt-4.1")
researcher = AssistantAgent(
"researcher",
model_client=model_client,
description="Looks up city facts with the lookup_city tool and drafts the answer.",
system_message=(
"You answer questions about cities. Use the lookup_city tool for every number "
"you report, and never guess a figure. Address the reviewer's feedback when given."
),
tools=[lookup_city],
reflect_on_tool_use=True,
max_tool_iterations=5,
model_context=BufferedChatCompletionContext(buffer_size=12),
)
reviewer = AssistantAgent(
"reviewer",
model_client=model_client,
description="Checks that every figure came from the tool and the answer is complete.",
system_message=(
"Check the draft. Every number must be traceable to a tool result, and the answer "
"must cover every city asked about. Give one concrete correction if it does not. "
"If the draft is correct and complete, reply with the single word APPROVE."
),
model_context=BufferedChatCompletionContext(buffer_size=12),
)
termination = TextMentionTermination("APPROVE", sources=["reviewer"]) | MaxMessageTermination(12)
team = RoundRobinGroupChat([researcher, reviewer], termination_condition=termination)
result = await Console(
team.run_stream(task="Compare Cairo, Riyadh and Dubai by population and timezone."),
output_stats=True,
)
print("\nstop_reason:", result.stop_reason)
print("usage:", model_client.total_usage())
with open("team_state.json", "w") as f:
json.dump(await team.save_state(), f)
await model_client.close()
asyncio.run(main())
Walk through the decisions, because each one is a thing you now know rather than a thing copied from a sample.
The tool returns JSON and, when it fails, returns a useful sentence rather than raising. That is deliberate: a model reads the string, so an error message listing the known cities lets it recover on its own turn. max_tool_iterations=5 is there because three cities need three lookups, and the default of one would have stopped after the first. reflect_on_tool_use=True is there because the answer should be prose comparing three cities, not three raw JSON blobs.
The termination condition has the sources=["reviewer"] restriction so the researcher cannot end the run by quoting APPROVE in a draft — a small, real failure mode. MaxMessageTermination(12) is the bound, and it counts the task message. Both agents get a buffered context so a long disagreement cannot grow unboundedly, and both description fields are written properly, which becomes necessary the moment you switch to SelectorGroupChat.
- Run the script. Confirm the reviewer approves and read the token stats.
- Remove a city from
CITY_DATAand run again. Watch the researcher handle the error string. - Write the second script that loads
team_state.jsonand asks "which of those three is furthest east?" without repeating the original task.
What you can now do, and what comes next
You can set up a pinned AutoGen environment and verify it without spending a token. You can create a model client for a hosted or a local model and describe an unrecognised model with model_info. You can build an agent with a system message and tools, and you know why max_tool_iterations and reflect_on_tool_use change its behaviour so much. You can put agents in a RoundRobinGroupChat, express "done" as a combined termination condition, and read stop_reason to find out what really happened. You can stream a run, count its tokens, bound its context, give it memory, save its state and resume it in another process. And you can read AutoGen's error messages and recognise a v0.2 or AG2 tutorial on sight.
What you have not touched is the rest of the framework. The next things to learn, roughly in order of how often they come up at work:
SelectorGroupChatandSwarm, for routing that is not a fixed order — includingselector_functo skip the selection model call when the choice is deterministic, andhandoffs=withHandoffTerminationfor explicit agent-to-agent transfer.- Workbenches and MCP, which is how an agent reaches real external systems. Remember
"mcp<2"and the "only connect to trusted MCP servers" warning. - Code execution with
CodeExecutorAgentandDockerCommandLineCodeExecutor, which became the default executor in 0.7.5 for good reason. Passapproval_func=or you will get a warning telling you code is running with no human oversight. The Docker guide covers the container side. - Structured output with
output_content_type=MyPydanticModel, which turns an agent's reply into a validated object instead of prose you have to parse. - Agents as tools —
AgentToolandTeamTool— remembering that they must not be called in parallel, so the parent's client needsparallel_tool_calls=False. GraphFlowandDiGraphBuilder, for fan-out, fan-in and loops, still marked experimental.- The Core API: runtimes,
AgentId, topics and subscriptions, andRoutedAgentwith@message_handler. This is where multi-tenancy and event-driven designs live, and it is the senior guide's territory. - Observability and deployment: OpenTelemetry spans through an OTLP backend, one team per session behind FastAPI, state in a database.
For context on where AutoGen sits, read two or three neighbours. CrewAI takes a role-and-task view of the same problem. LangGraph makes the state machine explicit from the start, which is closer to GraphFlow than to RoundRobinGroupChat. Semantic Kernel is the other parent of the successor framework, so time spent there is directly useful. The OpenAI Agents SDK is the lighter-weight single-vendor comparison. And once anything you build runs for someone other than you, Langfuse is how you find out what it actually did.
A closing word on judgement. Learn AutoGen for what it teaches: the vocabulary of agents, teams, tools, handoffs and termination is now shared across the field, and the concepts transfer almost one-for-one to Microsoft Agent Framework. If you own AutoGen code at work, pin it carefully, keep up with security fixes yourself, and plan a migration with the official guide in hand rather than under time pressure. If you are starting something new, start it on the successor — and this guide taught you most of what you need for that too.
Next in this series: the Mid-level guide, which picks up at SelectorGroupChat, workbenches, structured output, state in a real service and debugging beyond Console.
Sources
- AutoGen repository and README (maintenance-mode notice)
- AutoGen releases
- python-v0.7.5 release notes
- AutoGen documentation (stable)
- Installation
- Quickstart
- AgentChat tutorial
- Models
- Agents
- Teams
- Termination
- Human in the loop
- Managing state
- Memory
- Serializing components
- Migration guide (v0.2 to v0.4+)
- Core concepts: architecture
- Model clients
- Tools
- Command line code executors
- Extensions installation and extras
- AutoGen Studio
- API reference
- autogen-agentchat on PyPI
- autogen-ext on PyPI
- Migrating from AutoGen to Microsoft Agent Framework
- Microsoft Agent Framework