تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
Google ADKLLMsFrameworks & agents3 مستويات91 قسمًايغطّي Google ADK 2.10دليل بالإنجليزية

The Complete Google ADK Guide

Build, evaluate and deploy multi-agent systems with Google’s Agent Development Kit. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
15sections
24examples

This is part one of three. It covers everything you need to build a working agent with Google's Agent Development Kit, not a teaser. By the end you will have installed ADK, created an agent that answers questions, given it a tool it can call, watched every decision it made in a browser-based inspector, stored facts across turns, routed work through a graph of steps, and driven the whole thing from your own Python script. Mid-level and Senior take the same topics further; nothing here is wasted.

Each section ends with a Try it task. Do them as you go. Agents are one of those topics where reading produces a comfortable illusion of understanding that collapses the first time a model calls your function with arguments you did not expect.

A note on versions before we start, because it matters more here than with most tools. This guide is written against ADK 2.10.0 for Python, released in September 2026. ADK 2.0 arrived in May 2026 and replaced a large part of the 1.x execution model, so a great deal of the tutorial material you will find online — blog posts, conference talks, Stack Overflow answers, and the memory of any chatbot you ask — describes 1.x and will quietly mislead you. Wherever 1.x did something differently, this guide says so explicitly.

What ADK is, and the problem it solves

The Agent Development Kit is Google's open-source framework for building, testing and deploying agents: programs where a language model is allowed to decide what happens next. You give the model an instruction, a set of functions it is permitted to call, and a conversation; the model chooses whether to answer directly, call one of your functions, call another function with the first one's result, or hand the conversation to a different agent entirely. ADK is the machinery that runs that loop safely, records what happened, and lets you deploy the result.

To see why that machinery is worth having, picture building it yourself. You start with a single call to a model API: a prompt in, text out. That is easy. Then you want the model to be able to look something up, so you describe a function in the prompt and ask the model to reply with JSON when it wants to call it. Now you need to parse that JSON, validate it, call the real function, format the result, append it to the conversation, and call the model again. Then you need a cap on how many times that can loop, because a model that keeps asking for the same lookup will otherwise bill you forever. Then you want the conversation to survive a server restart, so you need storage. Then a file the user uploaded needs to be available three turns later, so you need somewhere to keep blobs. Then someone asks why the agent said something strange on Tuesday, and you discover you logged almost nothing useful.

Every one of those problems is generic. None of them is about your actual product. ADK's proposition is that it solves all of them, consistently, with components you can swap: the loop, the function-calling plumbing, the call cap, conversation storage, file storage, long-term memory, tracing, evaluation and deployment. What you write is the interesting part — the instructions, the tools, and the shape of the workflow.

Two design decisions shape everything that follows, and they are worth noticing now rather than discovering later.

ADK is code-first. An agent is an ordinary Python object you construct in an ordinary Python file. A tool is an ordinary Python function. There is no visual builder you must use and no domain-specific language you must learn, which means your agent is reviewable in a pull request, testable with pytest, and diffable in git log like anything else you write. A YAML form exists for configuring agents, but the normal path is code.

ADK is model-agnostic and deployment-agnostic. It is built by Google and optimised for Gemini models running on Google Cloud, and that is the smoothest path, but it is not a lock-in. ADK supports Anthropic's Claude models, OpenAI models, Oracle's OCI Generative AI service, and through LiteLLM a very long list of other providers including locally-hosted models served by Ollama. Likewise an ADK agent can run on your laptop, in a container anywhere, on Cloud Run, on Kubernetes, or on Google's fully managed Agent Runtime. For readers working with Gulf or Egyptian employers, that flexibility is not a detail: data-residency rules frequently decide which model endpoint and which region you are allowed to use, and a framework that treats the model as a swappable component is much easier to get past a security review than one that assumes a single provider.

Try it
  1. Think of one task at work that currently needs a person to look something up in two different systems and combine the answers.
  2. Write down the two lookups as function signatures, with argument names and return types.
  3. Write one sentence telling a new colleague when to use each lookup.
you now have the skeleton of an ADK agent: that sentence becomes the instruction, and those two signatures become the tools. The rest of this guide is mechanics.

The agent loop, drawn once

Before any code, fix the shape of the thing in your head. One user message produces one invocation, and an invocation is a loop.

USER MESSAGEone turn begins
→
MODELdecides: answer or call
→
TOOLyour Python runs
→
MODEL AGAINsees the result
→
FINAL ANSWERturn ends

The arrow from the tool back to the model is the part that makes this an agent rather than a chatbot. The model is not called once; it is called repeatedly until it decides it has enough to answer. Each pass around the loop costs a model call and therefore money and latency, and nothing in the model guarantees it will ever stop. That is why ADK caps the loop: by default an invocation may make at most 500 model calls, after which it raises an error rather than continuing. You will almost certainly never approach 500 legitimately. If you hit it, you have a bug, and the cap has just saved you from paying for it.

Everything the loop does is recorded as a stream of events. A user message is an event. A model response is an event. A request to call a tool is an event, and the tool's result is another. A change to stored state is an event. Events are immutable and ordered, which is what makes the whole system debuggable: when an agent behaves strangely, you do not guess, you read the event list for that invocation and see exactly which tool returned what.

Hold on to one more idea, which is where ADK 2.x differs most from what you may have read. Agents are not the only things that can be a step in the loop. In 2.x, an agent, a plain Python function, a tool and a whole nested pipeline are all nodes, and a pipeline is a graph of nodes with edges between them. The class BaseAgent is itself a subclass of BaseNode. This sounds abstract until you need a step that does no reasoning at all — reformat this JSON, write this row to the database — at which point being able to drop a plain function into the pipeline, with no model call and no cost, is obviously the right answer.

Try it
  1. Take the two-lookup task from the previous section.
  2. Count how many times around the loop a correct answer needs, assuming the model calls one tool at a time.
  3. Now imagine the model misreads your first tool's error message and retries it forever. Write down what you would want the framework to do.
three or four passes for a correct answer, and for the runaway case you will have described something very close to max_llm_calls. Good framework defaults usually encode a lesson someone learned expensively.

The nouns you need: agent, tool, session, runner

Four words carry most of the weight. Learn them precisely, because they are used loosely in conversation and precisely in the API.

Agent Tool Session Runner
Is A model plus an instruction plus permissions A capability the model may invoke One conversation thread The orchestrator
You write Agent(name=..., model=..., instruction=...) A Python function with type hints Rarely by hand Runner(app_name=..., agent=..., session_service=...)
Holds Reasoning and delegation Deterministic code state, events, ids Nothing; it drives the others
Analogy An employee with a job description A system they are allowed to use A ticket thread The shift supervisor

An agent is the model-driven unit. The class you will use almost always is LlmAgent, and Agent is simply an alias for it, so the two names mean the same thing. Its important fields are name (an identifier, used in traces and in delegation), model (a model identifier string), instruction (the system prompt that tells it what to do), description (one sentence saying what this agent is for, which matters because it is what other agents read when deciding whether to hand work over), and tools (the list of functions it may call). The default model in ADK 2.10 is gemini-3.5-flash, so you can leave model out, though being explicit is better practice.

A tool is, in the simplest and most common case, a plain Python function. You do not subclass anything or register anything. You write a function with type-annotated arguments and a docstring, pass it in the tools list, and ADK wraps it as a FunctionTool automatically, generating the schema the model needs from your annotations and the description from your docstring. That mechanism has a consequence worth saying out loud: your docstring is production code. It is the only explanation the model gets of what your function does and when to call it. A vague docstring produces a tool the model calls at the wrong moments, and a missing type hint produces arguments of the wrong type.

A session is one conversation thread. It is identified by three things together — app_name, user_id and a session id — and it holds the ordered list of events plus a key-value scratchpad called state. Sessions are created and loaded by a SessionService, which is the first component most people swap: InMemorySessionService for tests, a SQLite file for local development, a real database in production.

A runner is the orchestrator. You hand it a user message and it does everything else: loads the session, drives the loop, calls the model, executes tools, persists events, and yields those events back to you as they happen. It is constructed with a required keyword-only session_service, plus either an agent, a node, or an App object. App is the top-level container for a deployed agent and carries cross-cutting configuration such as plugins; you do not need it to get started.

Two smaller nouns will appear often enough to introduce now. An artifact is a named, versioned binary blob — a PDF, an image, a generated chart — stored by an ArtifactService rather than stuffed into state. Memory is long-term recall across sessions, served by a MemoryService, and is a different thing from state, which lives only within one conversation. Beginners conflate these constantly; the distinction is simply scope.

Why description is not optional in spirit instruction is read by this agent's model. description is read by other agents' models when they decide whether to delegate. If you ever build a multi-agent system and delegation goes to the wrong place, the first thing to check is whether your descriptions actually distinguish the agents from each other.
Try it
  1. Write the four nouns on paper from memory, with one sentence each.
  2. For each one, name the thing in a web application it most resembles.
  3. Say out loud the difference between state and memory.
if the state-versus-memory sentence comes out as "state is this conversation, memory is across conversations", you have it. That one is asked in interviews.

Installing ADK and checking the setup

ADK for Python requires Python 3.10 or later. Check before anything else, because the failure if you are on an older interpreter is a confusing dependency-resolution error rather than a clear message.

Always install into a virtual environment. Agent frameworks pull in a substantial dependency tree — the Google GenAI SDK, FastAPI, Uvicorn, Starlette, Pydantic, OpenTelemetry, SQLAlchemy's async pieces and more — and you do not want that mixed into your system Python.

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install google-adk

On Windows the activation step differs. In Command Prompt:

BAT
python -m venv .venv
.venv\Scripts\activate.bat
pip install google-adk

In PowerShell:

POWERSHELL
python -m venv .venv
.venv\Scripts\Activate.ps1
pip install google-adk

If PowerShell refuses to run the activation script, the fix is a one-off policy change: Set-ExecutionPolicy -Scope CurrentUser RemoteSigned. If you prefer uv, uv venv && uv pip install google-adk does the same job considerably faster.

Now verify, and verify in two ways, because the two checks fail differently and tell you different things.

BASH
adk --version
TEXT
adk, version 2.10.0
BASH
python -c "import google.adk; print(google.adk.__version__)"

The first proves the command-line tool is on your PATH. The second proves the library is importable from the interpreter you are actually using. If adk --version works but the import fails, you have activated one environment and installed into another — the single most common setup problem, and worth ten seconds of checking before you debug anything else. pip show google-adk is a third option and prints the version alongside the install location, which is useful precisely when you suspect that mismatch.

ADK ships a core install plus a long list of optional extras, and this is deliberate: the full set of integrations would be an enormous install for someone who wants one agent and one tool. The extras available in 2.10 include a2a, agent-identity, all, antigravity, benchmark, bigquery-analytics, community, daytona, db, dev, docs, e2b, eval, extensions, gcp, livekit, mcp, mongodb, oci, openai, otel-gcp, redis, slack, test, toolbox and tools. You install them with bracket syntax, and in zsh — which is the default shell on macOS — you must quote the brackets or the shell will try to glob them:

BASH
pip install "google-adk[eval,gcp]"

The three you are most likely to need early are db (database-backed sessions, which pulls in SQLAlchemy), eval (the evaluation tooling behind adk eval) and gcp (Vertex AI, Cloud Storage and Agent Runtime features). ADK is good about telling you when an extra is missing: reaching for a database session service without it produces the exact message The 'sqlalchemy' package is required to use this feature. Please install it by running: pip install google-adk[db], which is a refreshingly actionable error.

Pin the version ADK ships a minor release roughly every two weeks, and several recent minor releases contained real behaviour changes. Put google-adk==2.10.0 in your requirements file rather than a bare google-adk, and upgrade on purpose after reading the release notes. ADK also publishes constraints files (constraints-3.11.txt, constraints-3.12.txt) for fully reproducible installs.
Try it
  1. Create a virtual environment, activate it, and install google-adk.
  2. Run both verification checks above and confirm they report the same version.
  3. Deactivate the environment and run adk --version again.
after deactivating, the command should not be found. That is the shape of the mismatch problem, produced deliberately so you recognise it when it happens by accident.

Giving the agent credentials

ADK does not ship a model. It needs credentials for one, and there are two ways in. Choosing between them is the step where most first attempts stall, so be deliberate.

The quick path is Google AI Studio: you obtain an API key and set two environment variables in a file named .env inside your agent's folder, which the ADK command-line tools load automatically.

INI
GOOGLE_GENAI_USE_ENTERPRISE=0
GOOGLE_API_KEY=your-key-here

The production path is Vertex AI, also presented in the documentation as Gemini Enterprise, which authenticates with Google Cloud credentials instead of a key:

INI
GOOGLE_GENAI_USE_ENTERPRISE=1
GOOGLE_CLOUD_PROJECT=my-project
GOOGLE_CLOUD_LOCATION=us-central1

followed by one command to establish Application Default Credentials on your machine:

BASH
gcloud auth application-default login

Note the variable name carefully, because this is a trap laid by every tutorial written before mid-2026. The flag used to be called GOOGLE_GENAI_USE_VERTEXAI and took values like TRUE. Since ADK 2.3 the correct name is GOOGLE_GENAI_USE_ENTERPRISE. The old name still works through a deprecation fallback, but it emits the warning GOOGLE_GENAI_USE_VERTEXAI is deprecated, please use GOOGLE_GENAI_USE_ENTERPRISE instead. If you see that warning, rename the variable in your .env; do not set both.

The reason these two paths are not interchangeable is worth internalising. The flag tells the GenAI client which service to talk to, and the credential must match. An API key with the enterprise flag set to 1, or Cloud credentials with it set to 0, produces an authentication failure that looks like a bad key rather than like a misconfiguration — typically API key not valid, 400 INVALID_ARGUMENT or 403 PERMISSION_DENIED. When you see any of those, check the flag against the credential type first, before you regenerate a key you did not need to regenerate.

For the rest of this guide, either path works. Use AI Studio if you want to be running in two minutes; use Vertex if you are on a work machine where a Cloud project already exists and loose API keys would not survive review.

The .env file holds a secret ADK's own scaffolding warns about this: ⚠️ WARNING: Secrets (like GOOGLE_API_KEY) are stored in .env. It also writes a .gitignore for you. Do not defeat that. .env is a local-development convenience only; in production the key or the service-account identity belongs in Secret Manager or in workload identity, never in an image or a repository.
Try it
  1. Pick a path and create the matching .env in an empty folder.
  2. Deliberately set GOOGLE_GENAI_USE_ENTERPRISE to the wrong value for your credential.
  3. Come back and run the first agent in the next section, then read the error. Fix it and run again.
a permission or invalid-argument error that says nothing about the flag. Seeing it once on purpose saves you an hour of seeing it by accident.

Your first agent, step by step

ADK expects a specific folder layout, and getting it wrong is the most common first-run failure. Make a parent folder — call it agents — and let ADK create the agent inside it.

BASH
mkdir agents
cd agents
adk create my_agent

adk create asks which model and which backend you want, then writes four files: .env, .gitignore, __init__.py and agent.py. You can skip the prompts with flags: --model, --api_key, --project and --region. There is also a --type flag offering code or config; leave it at the default code, because the config variant is explicitly marked experimental and not ready for use.

The generated agent is exactly this:

my_agent/agent.py
from google.adk.agents.llm_agent import Agent

root_agent = Agent(
    model='gemini-3.5-flash',
    name='root_agent',
    description='A helpful assistant for user questions.',
    instruction='Answer user questions to the best of your knowledge',
)

and my_agent/__init__.py contains one line:

my_agent/__init__.py
from . import agent

Three things in that tiny file are load-bearing and will each bite you if changed carelessly. The variable must be called root_agent, because that is the name ADK's loader looks for. The __init__.py must import the agent module, or the loader cannot reach inside the package. And the agent lives in a folder whose name is the app name, inside a parent directory you run the tools from.

Now run it. The simplest form is an interactive terminal chat:

BASH
adk run my_agent

Note the working directory: you are in agents, the parent, and you pass the agent's folder name. For a single question with no chat loop, pass the query as a second argument:

BASH
adk run my_agent "what can you do?"

If instead you see this, the layout is wrong:

TEXT
No root_agent found for 'my_agent'. Searched in 'my_agent.agent.root_agent',
'my_agent.root_agent' and 'my_agent/root_agent.yaml'.

The message is unusually helpful: it lists the three places it looked. In order of likelihood, the cause is that you ran the command from inside my_agent instead of from agents, that __init__.py is empty or missing, or that you renamed the variable. A related error, A 'root_agent' was found for 'my_agent' but it is not a ... 'root_agent', means the name exists but holds something that is not an agent, node or App — commonly a function you forgot to call.

adk run has flags worth knowing on day one. --save_session writes the conversation to a file when you exit, optionally with --session_id ID; --resume FILE.json picks a saved session back up; --replay FILE.json re-runs a recorded set of inputs, which makes bug reports reproducible. --state '{"k":"v"}' seeds the session's state before the first turn. --in_memory skips local persistence entirely. --timeout 30s bounds a run. --log_level debug turns on the detail you need when the agent is doing something inexplicable, and --jsonl emits machine-readable output for scripting.

Try it
  1. Run adk create my_agent and then adk run my_agent, and have a short conversation.
  2. Exit, then cd my_agent and run adk run my_agent again. Read the error.
  3. Go back up, run with --save_session --session_id demo, exit, and resume it.
step two reproduces the No root_agent found error on demand. Step three shows you that a session is a real, portable object rather than something that only exists in the process.

Giving the agent a tool

An agent without tools can only talk. The moment it can call a function, it can act — and that is the whole point. Here is a complete agent with one tool.

my_agent/agent.py
from google.adk.agents.llm_agent import Agent

SHIPPING_DAYS = {"AE": 2, "SA": 3, "EG": 5, "KW": 3}

def estimate_delivery_days(country_code: str) -> dict:
    """Estimate delivery time in business days for a destination country.

    Use this whenever the user asks how long an order will take to arrive.

    Args:
        country_code: Two-letter ISO country code, uppercase, such as "AE" or "EG".
    """
    days = SHIPPING_DAYS.get(country_code.upper())
    if days is None:
        return {"status": "unsupported", "country_code": country_code}
    return {"status": "ok", "country_code": country_code.upper(), "days": days}

root_agent = Agent(
    model="gemini-3.5-flash",
    name="shipping_agent",
    description="Answers questions about order delivery times.",
    instruction=(
        "You help customers with delivery questions. "
        "When asked how long delivery takes, call estimate_delivery_days with the "
        "two-letter country code. If the tool reports status 'unsupported', say we "
        "do not ship there yet and do not guess a number."
    ),
    tools=[estimate_delivery_days],
)

Read that function again as the model sees it. The first docstring line is the summary that tells the model what the tool is for. The second line is an explicit trigger condition, which is far more effective than hoping the model infers one. The Args block explains the exact format expected, which is why the model sends "AE" rather than "United Arab Emirates". The country_code: str annotation becomes the parameter's type in the schema the model is given.

Four habits make tools behave well, and all four are visible above.

Return a dictionary, not a bare string. The model reads the result as data. A dict with a status key lets the model distinguish success from failure without parsing prose, and lets you add fields later without breaking anything.

Make failure a value, not an exception. estimate_delivery_days returns {"status": "unsupported"} instead of raising. The model can then say something sensible. An uncaught exception is for genuine bugs, not for "the user asked about a country we do not serve".

Say what to do with each outcome in the instruction. The line "do not guess a number" exists because a model with no number and a desire to be helpful will invent one. Instructions that cover the failure path are what separate a demo from something you would show a customer.

Keep one tool to one job. A tool that takes a mode argument and does three different things is three tools wearing a coat, and the model will pick the wrong mode.

One more detail about instructions, and it is a recent change. ADK substitutes session state into instructions using {braces}: if state holds customer_tier, then writing {customer_tier} in the instruction injects its value, and {customer_tier?} makes it optional so a missing key does not error. As of 2.10, ${var} and backslash-escaped \{var} are deliberately left exactly as written rather than substituted. If you are following an older tutorial that uses ${var} and wondering why nothing is injected, that is why.

Try it
  1. Put the shipping agent above into my_agent/agent.py and run it. Ask about Egypt, then about Japan.
  2. Now delete the Args block from the docstring and ask about Egypt again, phrasing the question with the country's full name.
  3. Restore the docstring and remove the "do not guess" sentence from the instruction. Ask about Japan.
step two often produces a tool call with the wrong argument format. Step three often produces an invented delivery time. The docstring and the instruction are not documentation; they are the program.

The dev UI, and reading what the agent actually did

Terminal chat is fine for a smoke test and useless for understanding behaviour. ADK bundles a local web interface that shows you the inside of every invocation.

BASH
adk web --port 8000

Run it from the parent directory — agents, not my_agent — because it presents a dropdown of every agent folder it can find. It binds to 127.0.0.1 on port 8000 by default. Open the address it prints, pick your agent, and talk to it.

The chat panel is the least interesting part. The panels that matter are Events and Trace. Events lists every event in the invocation in order: your message, the model's decision, the function call with its exact arguments, the function's exact return value, and the final response. When your agent does something baffling, the answer is almost always visible here within thirty seconds, and the usual revelation is that the model called your tool with an argument you had not considered. Trace shows the same invocation as timed spans, so you can see which model call or which tool consumed the latency. There are also panels for session state and for stored artifacts, and an Eval tab for running saved test cases.

Useful flags: --reload_agents picks up edits to your agent files without a restart, which turns the edit-instruction-and-retry cycle from twenty seconds into two. --logo-text and --logo-image-url let you rebrand it for a demo. --no-reload disables the file watcher, which you need on Windows in some configurations.

BASH
adk web --port 8000 --reload_agents

By default, local runs persist their state next to your agent: sessions go into <agents_dir>/<agent>/.adk/session.db, a SQLite file, and artifacts into <agents_dir>/<agent>/.adk/artifacts. Memory is in-process only. This is why a conversation can survive restarting adk web, and also why a stale .adk/session.db can cause schema errors after you upgrade ADK across a major version. For a local development database the fix is simply to delete the file and start again; for anything you care about, there is a proper migration command, covered below.

adk web is for your laptop only Its own help text says so: "This server is intended for local development. Its endpoints are unauthenticated." There is no login, no authorization, and no rate limiting. Do not expose it on a shared network, do not port-forward it from a cluster, and do not deploy it with --with_ui and call that production. The headless equivalent, adk api_server, carries the same warning and belongs behind your own authentication layer.
Try it
  1. Start adk web --reload_agents and ask the shipping agent about the UAE.
  2. Open the Events panel and find the exact arguments the model sent to your function.
  3. Edit the instruction in your editor, save, and ask again without restarting the server.
the Events panel shows the real call, not the one you imagined. This is the single habit that most shortens the time you spend confused by agents.

State, output_key, and remembering things

Agents need to remember. ADK gives you three distinct mechanisms, and using the wrong one is a classic beginner mistake, so separate them clearly.

State is a key-value scratchpad attached to a session. It is the right place for the small facts this conversation has established: the order number the user mentioned, the language they prefer, the result of step two. Keys carry an optional prefix that sets their scope, and the prefixes are the whole design:

Prefix Scope Use it for
none This session only Facts about this conversation
user: All sessions of this user in this app Preferences that should persist across conversations
app: All users of the app Shared configuration — never anything per-user
temp: This invocation only, never persisted Scratch values within one turn

The app: row deserves a warning sign. It is shared by every user of the application. Writing one user's data into an app: key is a data leak, not a bug in your logic, and it is the kind of mistake that is trivially easy to make and extremely awkward to explain afterwards.

The simplest way to write state is the output_key field on an agent. Set it, and the agent's final text is saved into state under that key automatically:

PYTHON
summariser = Agent(
    model="gemini-3.5-flash",
    name="summariser",
    description="Condenses a support ticket into two sentences.",
    instruction="Summarise the user's issue in at most two sentences.",
    output_key="ticket_summary",
)

A later agent can then read it in its own instruction with {ticket_summary}, and that pair — output_key writing, {braces} reading — is how most multi-step ADK pipelines pass data along without any plumbing code.

Tools write state too. A tool function can declare a ToolContext parameter, and ADK will supply it; the context exposes state for reading and writing, and save_artifact and load_artifact for files.

Artifacts are the second mechanism: named, versioned binary blobs — an uploaded PDF, a generated image — stored by an ArtifactService rather than embedded in state. Putting a base64 image into state technically works and is a bad idea, because state is loaded with every turn.

Memory is the third: recall that crosses session boundaries. It is served by a MemoryService — InMemoryMemoryService for development, Vertex AI Memory Bank or Vertex RAG in production — and the agent reaches it through the built-in load_memory and preload_memory tools. The distinction from state is purely one of scope: state is this conversation, memory is everything this user has ever told you.

For local development you get persistence for free, since sessions land in .adk/session.db. To be explicit, or to point somewhere else, the service URI flags work on run, web and api_server alike. --session_service_uri accepts memory://, sqlite://<path>, or a full SQLAlchemy URL. There is a rule about those URLs that catches everyone exactly once: ADK's database session service uses SQLAlchemy's async layer, so the URL needs an async driver. postgresql+asyncpg://... works; a bare postgresql://... does not, and produces Database related module not found for URL or Invalid database URL format or argument. The async spellings are sqlite+aiosqlite, postgresql+asyncpg, mysql+aiomysql and mariadb+asyncmy. You also need the db extra installed.

Try it
  1. Add output_key="last_answer" to your shipping agent and ask a question in adk web.
  2. Open the state panel and find the key.
  3. Add {last_answer?} into the instruction and ask a second question. Note the ?: on the first turn the key does not exist yet.
the second turn's instruction now contains the first turn's answer. Remove the ? and you will see why optional markers exist.

Workflows: graphs instead of fixed pipelines

Not every step should be decided by a model. When you know the order — classify, then route, then format — encoding that order in code is cheaper, faster and far more predictable than writing an instruction that begs a model to follow steps.

In ADK 2.x the construct for this is Workflow: a graph of nodes connected by edges. This is the biggest change from 1.x and the place where old tutorials will most reliably mislead you, so here is the honest version. ADK 1.x offered three orchestrator classes — SequentialAgent, ParallelAgent and LoopAgent — for running steps in order, in parallel, and in a loop. In 2.x all three still import and still run, but they emit a deprecation warning that says exactly what to do:

TEXT
DeprecationWarning: SequentialAgent is deprecated in favor of Workflow and will be
removed in a future version. Workflow cannot yet be used as an LlmAgent sub-agent.

Learn Workflow. Recognise the other three when you meet them in old code, and know they are on the way out.

A workflow is a Workflow object with a list of edges. Here is a complete, runnable triage graph:

PYTHON
from google.adk import Workflow, Event
from google.adk.workflow import DEFAULT_ROUTE

def classify(node_input: str):
    return Event(route="BUG" if "crash" in node_input else "OTHER", output=node_input)

def bug(node_input: str):
    return Event(message=f"bug handler got: {node_input}")

def other(node_input: str):
    return Event(message="other handler")

root_agent = Workflow(
    name="triage",
    edges=[
        ("START", classify),
        (classify, {"BUG": bug, DEFAULT_ROUTE: other}),
    ],
)

Four mechanics are on display, and they are the whole of the graph model.

START is the entry point. The edge ("START", classify) says the first node to run is classify. A tuple with more than two entries, such as ("START", a, b, c), is shorthand for a chain: a, then b, then c.

Data flows through node_input and output. The first node receives the user's text as its node_input argument. Whatever a node puts in its event's output becomes the next node's node_input. No shared mutable blackboard is required for the common case.

A route chooses the edge. When a node returns an Event with route="BUG", ADK looks at that node's outgoing edge dictionary and follows the "BUG" key. DEFAULT_ROUTE is the fallback for anything unmatched. A route may also be a list of strings, in which case execution fans out to several nodes at once.

Nodes are not only functions. An Agent can be a node, a nested Workflow can be a node, and so can a tool. The plain functions above cost nothing to run, which is the point: classify here is deterministic string matching, not a model call. Replace it with an Agent the day the classification genuinely needs judgement, and the graph around it does not change.

For fan-in, JoinNode waits for several predecessors to finish before continuing. max_concurrency on the workflow bounds how many nodes run in parallel. Individual nodes accept a timeout, which raises NodeTimeoutError when exceeded, and a retry_config built with RetryConfig(max_attempts=..., initial_delay=..., backoff_factor=...).

Two 2.x rules that are easy to break by accident A broad except Exception inside a node disables ADK's automatic retries, because the framework never learns the node failed. Let real exceptions propagate. And since 2.9, a node that failed runs again when an invocation resumes rather than replaying as if complete, so a node body with side effects must be idempotent. Both of these are silent in testing and loud in production.
Try it
  1. Save the triage workflow as your root_agent and run it with adk run, sending a message containing the word "crash" and then one without it.
  2. Change the classify function to return route="URGENT" for messages containing "urgent" and add a third handler.
  3. Delete the DEFAULT_ROUTE entry and send an unmatched message.
step three shows why a fallback edge is not optional in practice. Routing is exhaustive matching, and the default is your else.

Running an agent from your own code

The CLI is for development. In an application you construct the Runner yourself and consume its events. Here is the complete pattern.

run_agent.py
import asyncio
from google.adk import Runner
from google.adk.sessions import InMemorySessionService
from google.genai import types

from my_agent.agent import root_agent


async def main():
    runner = Runner(
        app_name="my_app",
        agent=root_agent,
        session_service=InMemorySessionService(),
    )
    session = await runner.session_service.create_session(
        app_name="my_app", user_id="u1"
    )

    message = types.Content(role="user", parts=[types.Part(text="How long to EG?")])
    async for event in runner.run_async(
        user_id="u1", session_id=session.id, new_message=message
    ):
        if event.is_final_response():
            print(event.content.parts[0].text)


asyncio.run(main())

Several details in those twenty lines are exactly the things people get wrong on a first attempt.

session_service is keyword-only and required. There is no default, and that is deliberate: where conversations are stored is an architectural decision, not something a framework should guess. Alongside it, the runner takes app_name plus one of agent, node or app. If you want the in-memory defaults with less typing, InMemoryRunner(agent=root_agent, app_name="my_app") wires them up for you — good for tests, never for production.

Create the session before you run, or opt out explicitly. Calling run_async with a session id that does not exist raises SessionNotFoundError with the message Session <id> not found.. If you would rather sessions appear on demand, pass auto_create_session=True to the Runner. Note that since 2.9 InMemorySessionService also raises this on appending to an unknown session, where it previously discarded the event silently — strictly better, and a change that surfaced some long-standing bugs in people's code.

The message must be a real types.Content with at least one part. Passing nothing produces ValueError: No parts in the new_message. or A new message is required for a new invocation..

run_async is an async generator and yields every event, not just the answer: partial streaming chunks, tool calls, tool results, state changes. That is a feature — it is how you build a UI that shows "looking up delivery times…" — but it means you must filter. event.is_final_response() is the public helper for "this is the answer". A synchronous runner.run(...) also exists for scripts, and since 2.8 it re-raises agent errors rather than swallowing them.

Try it
  1. Save the script above and run it against your shipping agent.
  2. Remove the is_final_response() filter and print event.author for every event instead.
  3. Delete the create_session call and run it again.
step two prints the full event stream and makes the loop from section two concrete. Step three produces SessionNotFoundError, which you now recognise on sight.

Configuration, and the errors you will actually hit

Most ADK configuration is environment variables and command-line flags rather than a config file, which is convenient once you know the names. A few flag conventions first: ADK's flags are snake_case with underscores — --session_service_uri, --log_level, --save_session — with a handful of hyphenated exceptions such as --reload/--no-reload, --provider-args, --logo-text and --allow-unsafe-unpickling. When a flag "does not exist", check the separator before anything else.

The environment variables worth knowing early:

Variable What it does
GOOGLE_GENAI_USE_ENTERPRISE Chooses Vertex/Gemini Enterprise (1) or AI Studio (0)
GOOGLE_API_KEY AI Studio key
GOOGLE_CLOUD_PROJECT, GOOGLE_CLOUD_LOCATION Vertex project and region
ADK_MAX_LLM_CALLS Overrides the 500-call-per-invocation cap

Now the errors. Here are the ones a beginner meets in their first week, with what each one actually means.

No root_agent found for '<name>'. Searched in '<name>.agent.root_agent', '<name>.root_agent' and '<name>/root_agent.yaml'. — the layout. You are inside the agent folder instead of its parent, or __init__.py does not import agent, or the variable is not called root_agent.

Agent not found: '<name>'. No matching directory or module ... — the app name in your URL or command does not match a folder. GET /list-apps on the API server lists the valid names. Since 2.5 an unknown app_name returns a sensible 404 rather than a 500.

Max number of llm calls limit of `500` exceeded (LlmCallsLimitExceededError) — the loop ran away. Two agents transferring to each other forever, or a tool whose error message makes the model retry indefinitely. Treat it as a bug to fix, not a limit to raise. If you genuinely need more headroom, RunConfig(max_llm_calls=N) or ADK_MAX_LLM_CALLS will give it, and a value of zero or less means unlimited with a warning — which you should essentially never want.

Session <id> not found. and Session with id <id> already exists. — the two halves of session lifecycle. Create it first, or auto_create_session=True; and do not call create_session with an id that is already taken when you meant get_session. One small gotcha from 2.10: the session id "user" is now reserved and rejected.

The 'sqlalchemy' package is required to use this feature. Please install it by running: pip install google-adk[db] — a missing extra, and the message tells you which. The same pattern applies for gcp, eval, a2a and the rest.

Model <name> not found. — you passed a bare model string that is not a Gemini model. Non-Google models need their wrapper class and their extra: LiteLlm("openai/gpt-...") from google.adk.models.lite_llm, or Claude/AnthropicLlm for Anthropic, or the OpenAI integration in google.adk.integrations.openai with pip install "google-adk[openai]".

google.genai.errors.ClientError: 429 RESOURCE_EXHAUSTED — model quota or rate limiting, not your code. Back off and retry; ADK has retry plugins and, since 2.9, automatic failover to a backup model.

A warning that the database is using the legacy v0 schema, or schema errors about missing node_info or output columns, means a session database created by an older ADK. The 2.0 event schema added fields. The command is adk migrate session --source_db_url ... --dest_db_url ..., and for a throwaway local database, deleting .adk/session.db is faster.

Try it
  1. Give your shipping agent an instruction that tells it to call estimate_delivery_days repeatedly until it is certain, and run it.
  2. Watch the Events panel, then stop it.
  3. Set ADK_MAX_LLM_CALLS=5 and run it again.
a loop, visible in the event list, and then a fast, clear failure. Lowering the cap during development turns runaway loops from an expensive mystery into a two-second error message.

Putting it all together

One small project that uses everything above. A support triage agent: it classifies an incoming message deterministically, routes technical issues to an agent with a diagnostic tool, sends everything else to a general agent, and stores the outcome in state.

triage_agent/agent.py
from google.adk import Workflow, Event
from google.adk.agents.llm_agent import Agent
from google.adk.workflow import DEFAULT_ROUTE

KNOWN_ISSUES = {
    "login": "Clear your browser cookies for the site and sign in again.",
    "upload": "Files must be under 25 MB. Larger uploads are rejected.",
}

def lookup_known_issue(topic: str) -> dict:
    """Look up the standard remedy for a known product issue.

    Call this before answering any technical question.

    Args:
        topic: A single lowercase keyword for the problem area, such as
            "login" or "upload".
    """
    remedy = KNOWN_ISSUES.get(topic.lower())
    if remedy is None:
        return {"status": "not_found", "topic": topic}
    return {"status": "ok", "topic": topic.lower(), "remedy": remedy}

def classify(node_input: str):
    """Deterministic triage. No model call, no cost."""
    text = node_input.lower()
    technical = any(word in text for word in ("error", "crash", "login", "upload"))
    return Event(route="TECHNICAL" if technical else "GENERAL", output=node_input)

technical_agent = Agent(
    model="gemini-3.5-flash",
    name="technical_agent",
    description="Diagnoses product errors using the known-issue list.",
    instruction=(
        "You are a support engineer. Always call lookup_known_issue first with a "
        "single keyword. If status is 'ok', give the remedy in your own words. "
        "If status is 'not_found', say the issue is not in the known list and ask "
        "for the exact error message. Never invent a remedy."
    ),
    tools=[lookup_known_issue],
    output_key="triage_outcome",
)

general_agent = Agent(
    model="gemini-3.5-flash",
    name="general_agent",
    description="Handles billing, account and general enquiries.",
    instruction=(
        "You handle non-technical support questions in at most three sentences. "
        "If the question is clearly a product fault, say it needs technical support."
    ),
    output_key="triage_outcome",
)

root_agent = Workflow(
    name="support_triage",
    edges=[
        ("START", classify),
        (classify, {"TECHNICAL": technical_agent, DEFAULT_ROUTE: general_agent}),
    ],
)

Walk through what happens when someone writes "I get an error when I upload a file". classify runs first as a plain function: it matches "error" and "upload", returns route="TECHNICAL", and passes the text through as output. ADK follows the "TECHNICAL" edge to technical_agent, which receives the text and, following its instruction, calls lookup_known_issue("upload"). The tool returns a dict with status: "ok" and a remedy. The model reads that and composes an answer. Because output_key is set, the final text is written to state["triage_outcome"], where any later step or any dashboard you build can read it.

Now notice what the structure buys you. The classification costs nothing, because it is string matching rather than a model call, and it is deterministic, so it is trivially unit-testable with plain pytest and no model at all. The expensive model calls happen only on the branch that needs judgement. The two agents have genuinely distinct description values, so if you later turn this into a delegating multi-agent system the model has something real to choose between. And the hallucination risk is addressed in the one place it can be: the instruction says explicitly what to do when the tool finds nothing.

Run it with adk web, send both kinds of message, and read the Events panel for each. On the technical branch you will see five events — your message, the classify node's routed event, the function call with topic: "upload", the function result, and the final response. On the general branch you will see three. That difference, visible in the event list, is the graph doing its job.

Try it
  1. Build the project above and send one message of each kind through adk web.
  2. Write a pytest test that calls classify("my upload crashed") and asserts the route is "TECHNICAL". No model, no network.
  3. Ask about something not in KNOWN_ISSUES and check that the agent asks for the error message instead of inventing a fix.
a passing test that runs in milliseconds, and an agent that admits ignorance. Both are consequences of pushing deterministic work out of the model and into code.

What you can now do, and what comes next

You can install ADK and verify it properly, configure either credential path and recognise the errors when they do not match, scaffold and run an agent, write tools whose docstrings actually steer the model, read an invocation's event list instead of guessing, persist facts in state at the right scope, build a graph that routes work between deterministic functions and model-driven agents, and drive the whole thing from your own async code. That is a genuinely useful agent, and you understand why each piece is there.

Several things were named and deliberately left shallow, and they are the natural next steps. Multi-agent delegation through sub_agents and transfer_to_agent, where the model chooses which specialist handles a request based on their descriptions. Callbacks — before_model_callback, before_tool_callback and their siblings — which let you inspect, modify or block a step, and are the main hook for guardrails. Plugins, which are bundles of callbacks applied globally to an App. Evaluation: eval sets, adk eval, and AgentEvaluator in pytest, which is how you stop agent quality from regressing silently when you change a prompt. MCP, the Model Context Protocol, which lets your agent consume tools from external servers through McpToolset and expose its own through to_mcp_server. And deployment: adk deploy cloud_run, adk deploy gke, adk deploy agent_engine for Google's managed Agent Runtime, and adk deploy docker for any container host.

Two warnings to carry forward. First, never run production behind in-memory services with more than one replica — sessions would be split across instances, and on Cloud Run or Kubernetes, where the agent directory is often read-only, ADK can silently fall back to in-memory storage unless you set service URIs explicitly. Second, the dev server is unauthenticated by design; real deployments need your own authentication layer, and the API server does not verify that the user_id in a request belongs to the caller.

For neighbouring tools in this catalogue, the obvious next reads are LangGraph for a different and instructive take on graph-based agent orchestration, Langfuse for tracing and evaluating agent runs across frameworks, MCP for the tool protocol ADK speaks, and Vertex AI for the managed platform ADK deploys onto most smoothly. If your agent will do retrieval, RAGAS covers measuring whether it retrieves well.

Then read Mid-level, which takes the execution model apart far enough that you can predict ADK's behaviour rather than observe it, and covers production services, testing, callbacks and the integration surface properly.

Sources