This is part one of three. It covers everything you need to build a working agent with the Claude Agent SDK, not a demo. By the end you can install the SDK, run an agent that reads and edits real files, decide which tools it may use, approve or refuse each action from your own code, keep a conversation going across turns, give the agent a tool you wrote, read the cost of a run, and recognise the error messages behind most first-week frustration. Mid-level and Senior take the same material into production hosting, sandboxing and cost control.
Each section ends with a Try it task. Do them as you go: an agent that calls tools stays abstract until you have watched your own process print Read, then Edit, then a result line with a dollar figure on it.
Everything here is verified against TypeScript SDK 0.3.285 and Python SDK 0.2.162, both released 2026-09-29 and both bundling Claude Code CLI v2.1.285. Examples lead with Python and give the TypeScript equivalent where the difference matters.
What the Claude Agent SDK is, and the problem it solves
The Claude Agent SDK is a library, published for Python and TypeScript, that runs Claude Code as a subprocess and exposes its agent loop to your program. Anthropic's own framing is the clearest available: build production AI agents with Claude Code as a library. You import a function, pass it a prompt and some options, and get back a stream of messages describing what the agent thought, which tools it called, what those returned, and what it concluded.
To see why that is useful, look at what you have to build without it. Suppose you want a program that reads a Python file, finds a bug, and fixes it. Using the raw Messages API — the subject of the Claude API guide — you would define a read_file and a write_file tool with JSON schemas, send the first message, inspect the response for tool-use blocks, execute each one yourself, append the results as a user message, and repeat until the model stops asking for tools. Then come the parts nobody warns you about: a loop guard so a confused model cannot spin forever, a diff-based edit tool because rewriting whole files is slow and lossy, a decision about what happens when the model tries to edit /etc/passwd, token accounting, and conversation persistence across restarts. That hand-written loop is several hundred lines before it is trustworthy, and it is the same several hundred lines in every project.
The SDK is that loop, already written and hardened, because it is the exact loop powering the Claude Code CLI. The tools come with it: Read, Edit, Write, Glob, Grep, Bash, WebSearch, WebFetch, Agent, Skill, ToolSearch and AskUserQuestion. So does the permission system, including an AST parser that inspects bash commands against your rules rather than matching strings. So does session persistence, as JSONL transcripts on disk, along with hooks, subagents, MCP servers, skills and plugins.
Two words in that diagram deserve attention, because almost every first-week surprise traces back to them.
Subprocess. When you call query(), the SDK launches an actual claude process and talks to it over standard input and output using newline-delimited JSON plus a small control protocol. One session means one subprocess. That process owns a shell, a working directory and a transcript file on local disk, and it never listens on a network port. The consequences are concrete: environment variables matter, because the subprocess inherits them or does not; the working directory matters, because that is where Read and Edit resolve relative paths; and if you want an agent behind an HTTP endpoint, you write that server yourself.
Claude Code. The agent you are driving is the same one people use interactively. That is a feature — it is unusually well-tested at reading and modifying codebases — but it also means the SDK inherits Claude Code's settings files, memory files and project configuration, which you will sometimes want to switch off.
It helps to place the SDK against its neighbours, because Anthropic ships four overlapping things. The Claude Code CLI is the interactive terminal interface: a human sits in front of it. The client SDK (anthropic, @anthropic-ai/sdk) makes direct Messages API calls where you own the tool loop; choose it for narrow, predictable model calls. Managed Agents put the agent loop on Anthropic's infrastructure, with sessions in an Anthropic-hosted or self-hosted sandbox, driven over REST, the SDKs or the ant CLI. The Agent SDK is for when you want the full coding-agent loop inside your own process, with your own code deciding what it may do. In any other language, the documented route is to run claude -p --output-format json as a subprocess yourself.
- Write down a task you would like an agent to do on a repository you own, in one sentence.
- List the tools it would need: reading files, searching, editing, running a test command.
- Mark which of those you would have to implement yourself with a raw API loop.
The agent loop, and the four nouns
Four nouns carry the whole model: turn, message, session and tool. Learn them in that order and the API stops looking like a wall of options.
A run begins when you hand the SDK a prompt. Claude responds with text, tool calls, or both. If there are tool calls, the SDK executes them — reading the file, running the bash command, calling your function — and feeds the results back. Claude responds again. The cycle repeats until Claude replies with no tool calls at all, which is how the loop knows the work is finished. Then you get a final assistant message and a result message.
One trip around that cycle is a turn. Memorise that precisely, because maxTurns counts tool-use turns only. Setting maxTurns: 5 on a task that genuinely needs twelve file reads does not make the agent efficient; it makes it stop halfway and hand you an error_max_turns result.
A message is one item in the stream you iterate over. There are five core types and you will use all of them.
SystemMessagecarries session metadata. Its subtypes includeinit, which arrives first and is the single most useful diagnostic in the SDK;compact_boundary, emitted when history is summarised; plusinformationalandworker_shutting_down.AssistantMessageis Claude speaking. Each message holds one content block, and blocks that came from the same API response share a message ID.UserMessagecarries tool results back into the conversation, along with anything you stream in yourself.StreamEventcarries raw token-level deltas, and only appears if you opt in withinclude_partial_messages.ResultMessageends the run. It carries the final text, tokenusage,total_cost_usd,num_turns,session_id, asubtype, astop_reasonand aterminal_reason.
Beyond those five there are observability messages — rate-limit events, task progress notifications, hook events, mirror errors, conversation resets. You need not handle them, but you must not crash on them, so always fall through gracefully on an unrecognised type.
A session is the persisted conversation, written as JSONL to ~/.claude/projects/<encoded-cwd>/<id>.jsonl, or under $CLAUDE_CONFIG_DIR/projects/. The encoded working directory is the real path with every non-alphanumeric character replaced by a hyphen. The critical point is that a session persists the conversation, not the filesystem. Resume tomorrow and Claude remembers that it edited utils.py; it does not restore utils.py. File restoration is a separate opt-in feature called checkpointing.
A tool is a capability the agent can invoke. The built-in set covers reading, editing, writing, globbing, grepping, running bash, searching and fetching the web, spawning subagents through the Agent tool, invoking skills, and asking the user a question. Read-only tools, and MCP tools that declare the readOnlyHint annotation, may run in parallel; mutating tools run one at a time. Custom tools default to sequential unless you mark them read-only.
system_prompt={"type": "preset", "preset": "claude_code"} in Python, or systemPrompt: { type: "preset", preset: "claude_code" } in TypeScript.
- Sketch the loop for "find every function longer than 50 lines and add a docstring", naming the tool used at each step.
- Count the turns. Then decide what you would set
max_turnsto. - Write down which of the five message types you would need to inspect to show progress in a UI.
Installing the SDK and checking the setup
You need Node.js 18 or later for TypeScript, or Python 3.10 or later for Python, and an API key from an Anthropic Console account at platform.claude.com. You do not need to install Claude Code separately in the normal case, because both SDKs ship a native Claude Code binary with them. There are exceptions, and they are the single most common installation failure, so they get their own treatment below.
For Python on macOS or Linux, work inside a virtual environment:
python3 -m venv .venv && source .venv/bin/activate
pip install claude-agent-sdk
Or with uv, which is faster and manages the environment for you:
uv init && uv add claude-agent-sdk && uv run agent.py
If you skip the virtual environment on a recent Debian, Ubuntu or Homebrew Python, pip refuses with error: externally-managed-environment. That message is not a problem with the SDK; it is your system Python protecting itself. Use a venv or uv.
For TypeScript, the full setup for a fresh project is four lines:
npm init -y && npm pkg set type=module
npm install @anthropic-ai/claude-agent-sdk
npm install --save-dev tsx
npx tsx agent.ts
Setting type=module matters because it enables top-level await, which keeps the examples short. In an existing CommonJS project, leave the package type alone and name the file agent.mts instead.
The native binary arrives through per-platform optional dependencies: -darwin-arm64, -darwin-x64, -linux-x64, -linux-arm64, -linux-x64-musl, -linux-arm64-musl, -win32-x64 and -win32-arm64. If your Docker build or CI pipeline runs npm ci --omit=optional, which many do to shrink images, no binary is installed and the SDK fails at run time with Native CLI binary for <platform>-<arch> not found. Keep optional dependencies, or set pathToClaudeCodeExecutable to a binary you installed another way. Yarn 1 installs both the glibc and the musl package; from SDK 0.2.141 onwards the correct one is selected at run time, and you may delete the unused one from an image to save space.
Python's wheel coverage has two gaps that matter where Alpine base images and Windows laptops are common. There is no musllinux wheel, so on Alpine you get the source distribution, which bundles no CLI at all. And although the documentation states that the Windows x64 wheel bundles claude.exe, versions 0.2.160, 0.2.161 and 0.2.162 ship no win_amd64 wheel — the 0.2.160 changelog notes CI skipped wheels over PyPI's per-file size limit. On Windows x64 today, pip install claude-agent-sdk resolves to the sdist with no binary.
In either case, install Claude Code natively and let the SDK find it on PATH, or point cli_path at it:
# macOS, Linux, WSL
curl -fsSL https://claude.ai/install.sh | bash
irm https://claude.ai/install.ps1 | iex
Homebrew (brew install --cask claude-code) and winget (winget install Anthropic.ClaudeCode) also work, though neither auto-updates. Pinning claude-agent-sdk==0.2.159 is the other Windows option. Check PyPI before you commit to either, because the wheel situation is a CI accident rather than a decision.
Authentication is one environment variable.
export ANTHROPIC_API_KEY=sk-ant-...
$env:ANTHROPIC_API_KEY = "sk-ant-..."
.env files. Nothing loads them for you. If your key lives in .env, load it yourself with python-dotenv or the dotenv package before you call query(), or you will get Invalid API key while staring at a file that plainly contains a valid key.
If your organisation routes model traffic through a provider, the SDK honours the Claude Code switches: CLAUDE_CODE_USE_BEDROCK=1 for Amazon Bedrock, CLAUDE_CODE_USE_VERTEX=1 for Google Cloud's Agent Platform, CLAUDE_CODE_USE_FOUNDRY=1 for Microsoft Foundry, and CLAUDE_CODE_USE_ANTHROPIC_AWS=1 with ANTHROPIC_AWS_WORKSPACE_ID for Claude Platform on AWS. Bedrock is what teams with data-residency requirements in the UAE and Saudi Arabia usually reach for, since the traffic stays inside an already-approved region; the AWS Bedrock guide covers that side. One policy point: unless Anthropic has approved it, third-party developers may not offer claude.ai login or subscription rate limits in their products. Use API-key authentication.
Verify in four steps. Confirm the package with pip show claude-agent-sdk or npm ls @anthropic-ai/claude-agent-sdk. Run a one-line agent and look for a result. Print the init system message, which lists the registered tools, every MCP server's status, the model, and apiKeySource — whose value will be ANTHROPIC_API_KEY, apiKeyHelper, /login managed key or, tellingly, none. And if you depend on a native install, run claude --version in the same environment as the service: IDE terminals, systemd units and container entrypoints routinely have a different PATH from your login shell.
- Install the SDK in a fresh virtual environment and export your API key.
- Print the
initmessage and readapiKeySourceand the tool list aloud. - If you are on Alpine or Windows, check whether a wheel with a binary was actually installed.
init message naming your model and about a dozen tools. If apiKeySource is none, fix that before writing another line.
Your first agent, step by step
Create a directory with one deliberately broken file in it, because an agent fixing a real bug teaches more than an agent summarising text.
def average(numbers):
total = 0
for n in numbers:
total += n
return total / len(numbers)
That function raises ZeroDivisionError on an empty list. Now the agent:
import asyncio
from claude_agent_sdk import (
query,
ClaudeAgentOptions,
AssistantMessage,
TextBlock,
ToolUseBlock,
ResultMessage,
)
async def main():
options = ClaudeAgentOptions(
allowed_tools=["Read", "Edit", "Glob"],
permission_mode="acceptEdits",
)
async for message in query(
prompt="Review utils.py for bugs that would cause crashes. Fix any issues you find.",
options=options,
):
if isinstance(message, AssistantMessage):
for block in message.content:
if isinstance(block, TextBlock):
print(block.text)
elif isinstance(block, ToolUseBlock):
print(f"[tool] {block.name}")
elif isinstance(message, ResultMessage):
print(f"Done: {message.subtype}")
if message.total_cost_usd is not None:
print(f"Cost: ${message.total_cost_usd:.4f}")
asyncio.run(main())
Run it with python agent.py. The output interleaves Claude's reasoning with tool names, and ends with something close to:
[tool] Glob
[tool] Read
[tool] Edit
I found a crash in average(): dividing by len(numbers) raises
ZeroDivisionError when the list is empty. I added a guard that
returns 0.0 for an empty input.
Done: success
Cost: $0.0231
Open utils.py. It has changed on disk. That is the moment the SDK stops being abstract.
Four things in those twenty lines deserve a sentence each. The function is async, and query() is an async iterator, because the agent loop is a stream of events arriving over time; there is no synchronous version. allowed_tools names three tools, and the significance of that list is subtler than it looks. permission_mode="acceptEdits" is what allows the edit to happen without a human approving it. And the final block checks total_cost_usd is not None before formatting it, because in Python total_cost_usd, usage and model_usage are all Optional.
The TypeScript shape is the same idea with different spelling. TypeScript uses camelCase throughout, discriminates messages on a type string rather than with isinstance, and places assistant and user content at message.message.content — a nesting that trips up everyone exactly once:
import { query } from "@anthropic-ai/claude-agent-sdk";
for await (const m of query({
prompt: "Summarize the unfinished work marked in this repository",
options: {
model: "claude-sonnet-5",
allowedTools: ["Read", "Glob", "Grep"],
maxTurns: 8,
cwd: "/path/to/repo",
},
})) {
if (m.type === "result" && m.subtype === "success" && !m.is_error) {
console.log(m.result);
}
}
Note the cwd option. It sets the working directory of the subprocess, which is what Read, Glob and Grep resolve paths against. Without it the agent works relative to wherever your program happens to have been started, which in a service is rarely where you think.
- Run the Python agent above against your own broken file and read the diff it produced.
- Remove
permission_modeentirely and run it again. Watch the edit get refused. - Add
"Bash"toallowed_toolsand ask the agent to run your tests after fixing the bug.
Reading the messages that come back
Your first agent printed text and tool names. A real program needs to know whether the run succeeded, and that is less obvious than it sounds.
The result message carries a subtype, and it is one of five values: success, error_max_turns, error_max_budget_usd, error_during_execution, or error_max_structured_output_retries. The final text lives in the result field, and that field only exists on success. Reading it unconditionally is the most common bug in beginner code, because it works perfectly until the day a run hits its turn cap and your handler raises AttributeError on a None.
There is a second, sharper rule that the documentation is explicit about and almost nobody discovers on their own: check terminal_reason before subtype. When the final request to the API fails, Claude Code can report subtype: "success" with terminal_reason: "api_error". If you trust subtype alone you will record a successful run that produced nothing. Alongside those, stop_reason tells you why the model stopped generating, and is typically end_turn, max_tokens or refusal.
from claude_agent_sdk import ResultMessage
def summarise(message: ResultMessage) -> str:
if message.terminal_reason and message.terminal_reason != "success":
return f"failed: terminal_reason={message.terminal_reason}"
if message.subtype != "success":
return f"failed: {message.subtype} after {message.num_turns} turns"
return message.result or "(no text)"
Every result also carries a session_id, whether or not the run succeeded. Log it. It is the handle you need to resume the conversation, to find the JSONL transcript on disk, and to correlate the run with traces if you export telemetry.
The cost fields need care about meaning rather than precision. total_cost_usd and the per-model costUSD figures are client-side estimates from a price table bundled with the CLI. They are good for dashboards and for spotting a runaway loop; they are not billing data, and you must never invoice from them — the Usage and Cost API and the Console are authoritative. The fields count different things: usage covers the main loop only and excludes subagents, while total_cost_usd and model_usage include them. From CLI 2.1.277 a resumed or forked session's total_cost_usd continues from earlier turns rather than restarting at zero, so read the latest result instead of summing.
Two failure shapes are worth rehearsing. When a single-shot query() ends on an error result, the SDK yields that result and then raises, so a loop without a try/except surfaces the exception rather than your message. And if the process crashes mid-run the final result is error_during_execution, possibly with zeroed cost; treat missing cost as unknown rather than free.
Python's error hierarchy was tightened in 0.2.140. A run that ends on an error result now raises ResultError, a subclass of ProcessError, carrying structured fields. Because it is a subclass, the except ResultError clause must come before except ProcessError, or the specific handler never runs:
from claude_agent_sdk import query, ResultError, ProcessError, CLINotFoundError
try:
async for message in query(prompt=task, options=options):
handle(message)
except ResultError as exc:
print("agent finished on an error result:", exc)
except ProcessError as exc:
print("the CLI process itself failed:", exc)
except CLINotFoundError as exc:
print("no claude binary:", exc)
- Set
max_turns=1on a task that needs several and read the full result object. - Wrap the loop in the try/except above and confirm which clause fires.
- Print
session_id, then find the matching JSONL file under~/.claude/projects/.
error_max_turns result with no result text, a ResultError, and a transcript file you can open in a text editor.
Choosing tools, and what allowed_tools really means
This section exists because one option name misleads almost every newcomer, and the misunderstanding is a security problem rather than a style problem.
allowed_tools is an auto-approve list. It does not restrict which tools exist. Naming ["Read", "Edit", "Glob"] means those three run without asking. It does not remove Bash from the agent's awareness. Bash is still registered, Claude can still call it, and what happens then depends on your permission mode — in default mode it falls through to your approval callback and is denied if there is none, but in bypassPermissions it simply runs.
Three different options shape the tool surface, and they do different jobs.
tools sets the base set. Pass ["Read", "Edit", "Bash"] for exactly those; pass [] for no built-ins at all, which is the right choice for an agent that should only use your own custom tools; or pass {"type": "preset", "preset": "claude_code"} for the full default set.
disallowed_tools removes capability. A bare name takes the tool out of context entirely — the model never learns it exists. A scoped rule such as Bash(rm *) denies matching calls in every permission mode, including bypassPermissions, which makes it the right tool for hard prohibitions.
allowed_tools grants silent approval to specific names or scoped patterns.
Two sensible starting shapes follow from that. A read-only analyst agent, safe to point at any repository:
options = ClaudeAgentOptions(
tools=["Read", "Glob", "Grep"],
allowed_tools=["Read", "Glob", "Grep"],
permission_mode="dontAsk",
)
And a locked-down worker that may edit inside one directory and nothing else:
options = ClaudeAgentOptions(
cwd="/srv/workspace",
allowed_tools=["Read", "Glob", "Grep", "Edit", "Write"],
disallowed_tools=["Bash", "WebFetch"],
permission_mode="acceptEdits",
max_turns=20,
)
Two recent defaults will confuse you if nobody mentions them. From TS 0.3.162, on native builds SDK sessions default to embedded find and grep inside Bash rather than registering dedicated Grep and Glob tools; name them explicitly in tools or allowed_tools if you want them as real tools a hook can inspect. And the task-tracking tools (TaskCreate, TaskGet, TaskUpdate, TaskList, plus the older TodoWrite) are default tools only on Claude 3.x, Opus 4.0 through 4.7, Sonnet 4.0 through 4.6 and Haiku 4.5. On Opus 4.8, Sonnet 5, Fable 5, Mythos 5 and newer you must name them yourself or set CLAUDE_CODE_ENABLE_TODO_TOOLS=1. TodoWrite was itself replaced by the Task tools in SDK sessions in TS 0.3.142, and the Task tools accumulate by ID rather than replacing a list.
The model option takes an alias or a full ID; the documentation uses claude-sonnet-5, claude-opus-5 and claude-fable-5. Leave it unset and the session uses Claude Code's default. Pass a display name like "Sonnet5" and you get a pointed error: Model "Sonnet5" is not a recognized model id. Did you mean 'claude-sonnet-5'?
Finally, give every agent a budget. Both caps are off by default, which means an agent in a loop will spend until something else stops it. max_turns bounds tool-use turns, and max_turns=0 means unlimited. max_budget_usd bounds estimated spend, and 0 is rejected at startup rather than treated as unlimited.
- Build the read-only agent above and ask it to write a file. Read the refusal.
- Add
disallowed_tools=["Bash"]and ask it to runls. Note that it does not even try. - Set
max_budget_usd=0.05on a large task and watch forerror_max_budget_usd.
Permissions without a human in the terminal
In the CLI a permission prompt is a dialog. In the SDK there is nobody to ask, so you get two mechanisms: a mode that sets the default posture, and a callback that decides case by case. Every tool call is evaluated in a fixed order, and the outcome is binary — allowed or denied:
- Hooks run first and can decide outright.
- Deny rules are checked next.
- Ask rules force a prompt even for otherwise-approved calls.
- The permission mode applies its default posture.
- Allow rules auto-approve what is left.
- The
can_use_toolcallback handles anything still undecided.
The modes, in rough order of how much they let through:
default— unapproved calls go tocan_use_tool. With no callback, they are denied. This is the safe starting point.acceptEdits— auto-approves file edits and the shell commandsmkdir,touch,rm,rmdir,mv,cpandsedinside the working directory or additional directories.plan— edits are never auto-approved, which suits an agent that should propose rather than act. As of TS 0.3.269, writes in plan mode route throughcan_use_tooleven whenallowDangerouslySkipPermissionsis set.dontAsk— anything that would prompt is denied outright andcan_use_toolis never called. For headless agents this is the recommended posture, paired with an explicitallowed_toolslist.auto— a model classifier approves or denies.bypassPermissions— everything is approved except explicit ask rules, deny rules, hooks, user-interaction tools and critical-pathrm/rmdir. TypeScript additionally requiresallowDangerouslySkipPermissions: true. It refuses to run as root or undersudooutside a recognised sandbox on Linux and macOS.
The alias 'manual' has been accepted for 'default' since TS 0.3.200, so do not be surprised to see it in sample code.
The callback is where interesting policy lives. In Python it receives the tool name, the input, and a context, and returns either an allow or a deny:
from claude_agent_sdk import (
ClaudeAgentOptions,
PermissionResultAllow,
PermissionResultDeny,
)
SAFE_PREFIXES = ("pytest", "ruff", "git status", "git diff")
async def gate(tool_name, input_data, context):
if tool_name == "Bash":
command = input_data.get("command", "")
if not command.startswith(SAFE_PREFIXES):
return PermissionResultDeny(message=f"command not on the allowlist: {command}")
if tool_name in ("Write", "Edit"):
path = input_data.get("file_path", "")
if path.endswith((".env", ".pem")):
return PermissionResultDeny(message="secrets are off limits")
return PermissionResultAllow()
options = ClaudeAgentOptions(
permission_mode="default",
can_use_tool=gate,
allowed_tools=["Read", "Glob", "Grep"],
)
PermissionResultAllow can also carry updated_input, which rewrites the call before it executes — clamping a path, adding --dry-run — and updated_permissions to adjust later decisions. PermissionResultDeny takes a message that goes back to the model, so write it as instruction rather than as a log line: "secrets are off limits, use the config loader instead" steers the next attempt where "denied" does not. It also accepts interrupt=True to stop the run.
TypeScript's equivalent is canUseTool(toolName, input, { signal, suggestions, ... }) returning { behavior: "allow", updatedInput } or { behavior: "deny", message }.
bypassPermissions already approved. Set can_use_tool together with either of those and you have a policy that silently does nothing — TypeScript warns with the code CLAUDE_SDK_CAN_USE_TOOL_SHADOWED, Python emits a runtime warning. If you need to see every call, use a PreToolUse hook instead. The one exception is AskUserQuestion, which always reaches the callback.
Hooks are the other half of the picture, and worth meeting at beginner level even if you use them later. A hook is a callback in your own process, fired at a lifecycle event — PreToolUse, PostToolUse, Stop, UserPromptSubmit, PreCompact and others — that can deny a call, modify its input, or add context. Hooks run outside the model's context, so they cost no tokens, and all hooks matching an event fire concurrently. Matchers are exact strings or regexes on the tool name:
from claude_agent_sdk import ClaudeAgentOptions, HookMatcher
async def protect_env_files(input_data, tool_use_id, context):
if input_data["tool_input"].get("file_path", "").endswith(".env"):
return {
"hookSpecificOutput": {
"hookEventName": input_data["hook_event_name"],
"permissionDecision": "deny",
"permissionDecisionReason": "Cannot modify .env files",
}
}
return {}
options = ClaudeAgentOptions(
hooks={"PreToolUse": [HookMatcher(matcher="Write|Edit", hooks=[protect_env_files])]}
)
Hook timeouts are set per matcher in seconds, and the defaults are 600 seconds for most events, 30 for UserPromptSubmit and the model-switch hooks, 10 for MessageDisplay and about 1.5 for SessionEnd. A hook that times out is discarded and the session continues, which is forgiving but means a slow hook can fail open.
- Add the
gatecallback and ask the agent to runcurl. Read the model's reaction to your deny message. - Set
allowed_tools=["Bash"]alongside the callback and watch the gate stop firing. - Replace it with the
PreToolUsehook and confirm the call is intercepted again.
Sessions: continuing, resuming and forking
By default each query() call in Python starts a fresh session with no memory of the last one. Three options change that.
Continue picks up the most recent session in the current working directory — continue_conversation=True in Python, continue: true in TypeScript. It is the right choice for a CLI-like tool where "the last conversation" is unambiguous.
Resume reopens a specific session by ID: resume="<session-id>". Because every result message carries session_id, the pattern is to store that ID with whatever your application calls a conversation, then resume it later. This is how you build a chat interface over the SDK.
Fork creates a new session ID starting from a copy of an existing history, leaving the original untouched: fork_session=True alongside resume. Forking is how you explore two approaches from the same context, or offer an "edit this message and retry" affordance without destroying the original thread.
first = None
async for message in query(prompt="Map the modules in this repo", options=options):
if isinstance(message, ResultMessage):
first = message.session_id
async for message in query(
prompt="Now write a README section describing them",
options=ClaudeAgentOptions(resume=first, allowed_tools=["Read", "Glob", "Write"]),
):
...
You can also supply your own session_id, but it must be a valid UUID, and it cannot be combined with continue or resume unless you also fork. Both SDKs ship session management helpers: list_sessions, get_session_messages, get_session_info, rename_session, tag_session, delete_session, fork_session, list_subagents and get_subagent_messages in Python, with camelCase equivalents in TypeScript. Naming and tagging sessions sounds cosmetic until you have a few hundred transcripts and need to find the one from the incident last Tuesday.
If you would rather nothing touched the disk, TypeScript offers persistSession: false to keep a session in memory only. Python has no such option; set CLAUDE_CODE_SKIP_PROMPT_HISTORY in env instead.
Two related features round this out. Compaction happens automatically as the conversation approaches the context limit: older history is summarised and a compact_boundary system message is emitted. You can trigger it by sending /compact as a prompt, and hook it with PreCompact (whose trigger is manual or auto). The practical lesson is where to put rules you need honoured all session: not in the first prompt, which may be summarised away, but in CLAUDE.md, which is re-injected on every request. File checkpointing is the answer to "my agent edited the wrong thing": enable enableFileCheckpointing and call rewindFiles(userMessageId) to restore files. It must be enabled on the original session and on any resumed session.
- Run two queries where the second resumes the first, and ask it "what did you just look at?".
- Fork that session, take the fork in a different direction, then resume the original and confirm it is unchanged.
- Put a rule in
CLAUDE.md— "never edit files undermigrations/" — and check it survives a long run.
Streaming input and the long-lived client
So far every prompt has been a string. That is single message input: one shot, stateless, and the right shape for a Lambda handler or a CI job, with real limitations — no image attachments, no queueing mid-run, no real-time interrupts.
Streaming input is the mode the documentation recommends for everything else. You pass an AsyncIterable of message dictionaries instead of a string, and in exchange you get a long-lived session: messages can be queued while the agent works, images attached, the run interrupted, the model and permission mode changed mid-session, and permission prompts answered inside the loop.
In Python the ergonomic way to get all of that is ClaudeSDKClient, which keeps one session alive across calls:
import asyncio
from claude_agent_sdk import ClaudeSDKClient, ClaudeAgentOptions, ResultMessage
async def main():
options = ClaudeAgentOptions(allowed_tools=["Read", "Glob", "Grep"])
async with ClaudeSDKClient(options=options) as client:
await client.query("Which module handles authentication?")
async for message in client.receive_response():
if isinstance(message, ResultMessage):
print(message.result)
await client.query("Now list every function it exports.")
async for message in client.receive_response():
if isinstance(message, ResultMessage):
print(message.result)
asyncio.run(main())
The async with block is not decoration. Calling a client method before connect() or after disconnect() raises Not connected. Call connect() first., and the context manager is how you never see it. receive_response() stops after the result message, which is what you want per turn; receive_messages() streams everything indefinitely, for a live view.
The client also exposes control methods that single-shot query() cannot offer: interrupt() to stop the agent mid-work, set_permission_mode() to tighten or loosen policy on the fly, set_model() (pass None to return to the default), rewind_files(id), get_mcp_status(), reconnect_mcp_server(), toggle_mcp_server(name, enabled), stop_task(id), get_server_info(), get_context_usage() and disconnect().
TypeScript has no client class. The Query object returned by query() carries the control methods instead — interrupt(), setPermissionMode(), setModel(), streamInput(), initializationResult(), supportedModels(), getContextUsage(), readFile(), reloadSkills() and more — and you keep continuity across calls with continue: true or resume. Four of those (interrupt, setPermissionMode, setModel, applyFlagSettings) require streaming input mode.
Claude Code process aborted by user, sometimes after a long minified line. In Python the exception is logged only at debug level and the session simply hangs. Wrap the body of your generator in a try/except that logs loudly.
- Convert your agent to
ClaudeSDKClientand ask three follow-up questions in one session. - Call
get_context_usage()after each turn and watch the number grow. - Start a long task, then call
interrupt()from a second coroutine.
Giving the agent a tool you wrote
Built-in tools cover files and shell. Everything specific to your business — a customer lookup, a metrics query, a deployment trigger — you supply as a custom tool. The SDK implements these as an in-process MCP server: you write a Python function and the plumbing is handled. The MCP guide covers the protocol; here you need two helpers.
import asyncio
from claude_agent_sdk import (
tool,
create_sdk_mcp_server,
query,
ClaudeAgentOptions,
ResultMessage,
)
@tool(
"get_temperature",
"Get the current temperature at a location",
{"latitude": float, "longitude": float},
)
async def get_temperature(args):
return {"content": [{"type": "text", "text": "Temperature: 61°F"}]}
server = create_sdk_mcp_server(
name="weather",
version="1.0.0",
tools=[get_temperature],
)
async def main():
options = ClaudeAgentOptions(
mcp_servers={"weather": server},
allowed_tools=["mcp__weather__get_temperature"],
)
async for message in query(prompt="How warm is it in Cairo?", options=options):
if isinstance(message, ResultMessage):
print(message.result)
asyncio.run(main())
The naming rule is the part to memorise, because getting it wrong produces a silent failure rather than an error. Tool names are fully qualified as mcp__{server_key}__{tool_name}, where the server key is the key in the mcp_servers dictionary — not the name= passed to create_sdk_mcp_server. Here that gives mcp__weather__get_temperature. If the name is wrong, the tool appears in the agent's list, Claude calls it, and the permission flow denies it: the symptom is "my MCP tools are visible but never run". Allow either the exact name or the whole server with mcp__weather__*. An unanchored * or a bare mcp__* in an allow rule is ignored with a warning, although deny rules do accept those globs.
TypeScript's version is tool(name, description, zodSchema, handler, { annotations }) with createSdkMcpServer({ name, version, tools, timeout }), using Zod for the schema instead of a dictionary of types. Python supports typing.Annotated for per-parameter descriptions, which is worth using because the description is what the model reads when deciding whether the tool fits.
Four annotations shape how a tool is scheduled: readOnlyHint, destructiveHint, idempotentHint and openWorldHint. The first has a concrete effect — custom tools run sequentially by default, and marking one read-only lets it run in parallel. To signal failure from inside a tool, return is_error in the result rather than raising.
You can also connect external MCP servers to reach a database or a hosted service someone else maintains. Configure them in the same mcp_servers mapping: stdio as {command, args, env}, remote as {type: "http" | "sse", url, headers}, or a path to a config file. Since TS 0.3.142 they connect in the background by default, so your session starts immediately and a slow server shows status: "pending" in init. Pending is not a failure — the statuses are pending, connected, failed, needs-auth and disabled. Set MCP_CONNECTION_NONBLOCKING=0 to restore blocking startup, capped at five seconds, or mark a server alwaysLoad: true.
- Write a custom tool that returns a hard-coded value and get the agent to call it.
- Deliberately misspell the allowlist entry and confirm the call is denied rather than erroring.
- Add
readOnlyHintand ask a question that needs two lookups at once.
mcp__server__tool convention burned in permanently by step two.
Structured output, for when you need JSON
Parsing prose with regular expressions is a habit the SDK lets you avoid. Ask for a JSON schema and you get a validated object.
schema = {
"type": "object",
"properties": {
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"files_changed": {"type": "array", "items": {"type": "string"}},
"summary": {"type": "string"},
},
"required": ["severity", "files_changed", "summary"],
}
options = ClaudeAgentOptions(
allowed_tools=["Read", "Glob", "Grep"],
output_format={"type": "json_schema", "schema": schema},
)
async for message in query(prompt="Audit this repo for crash bugs", options=options):
if isinstance(message, ResultMessage):
if message.subtype == "success" and message.structured_output is not None:
print(message.structured_output["severity"])
else:
print("structured output unavailable:", message.subtype)
TypeScript uses outputFormat: { type: "json_schema", schema }. In both languages you can generate the schema rather than hand-writing it — Pydantic in Python, Zod in TypeScript — which keeps one definition for validation and for the model.
The rule that saves you an incident: require both subtype == "success" and a non-null structured_output. A schema the model cannot satisfy produces a successful-looking result with no structured output, and there is a dedicated subtype, error_max_structured_output_retries, for repeated failures. Treat a missing object as a failure and simplify the schema — flat beats deeply nested. A related error, API Error: 400 ... tools.N.custom.input_schema: JSON schema is invalid, means a tool schema is malformed; schemas must be valid draft 2020-12 with property keys matching ^[a-zA-Z0-9_.-]{1,64}$.
- Define a three-field schema and run the audit agent against a real repository.
- Add a required field the repo cannot support and observe the empty structured output.
- Regenerate the schema from a Pydantic model instead of writing it by hand.
Configuration, settings sources and the errors you will actually see
Two configuration surfaces cause more beginner confusion than all the others combined, and both concern things the SDK reads from your filesystem or environment without being asked.
Setting sources. Omitting setting_sources loads user, project and local filesystem settings exactly as the CLI does: ~/.claude/settings.json, .claude/settings.json, .claude/settings.local.json, CLAUDE.md, and the skills, agents and commands under .claude/. That is equivalent to ["user", "project", "local"]. Pass [] to isolate. The default has a confusing history — the 0.1.0 changelog announced "no filesystem settings by default", but that was reverted and the line is stale, so verify with the init message rather than trusting either document. Some sources are read regardless of the option: managed policy, ~/.claude.json, auto memory, claude.ai connectors and sandbox credential denies. So setting_sources=[] is strong isolation but not complete, which is why the Senior guide spends real time on multi-tenancy.
Two symptoms follow from this option. If skills are not found, either setting_sources excludes user or project, or your cwd is outside the repository holding .claude/skills/. And if your agent behaves differently on a colleague's machine, compare ~/.claude/settings.json and CLAUDE.md first.
Environment variables. The env option differs between the two SDKs in a way that is responsible for a remarkable share of first-day failures. In TypeScript, env replaces process.env for the subprocess. Set env: { MY_VAR: "x" } and you have just removed ANTHROPIC_API_KEY, PATH and everything else from the agent's environment, and you will get Not logged in · Please run /login or Invalid API key. Always spread:
const options = { env: { ...process.env, MY_VAR: "x" } };
In Python, env is merged over the inherited environment, so env={"MY_VAR": "x"} behaves as you would expect. The asymmetry is deliberate and documented; it is simply surprising.
A handful of other beginner-relevant options are worth knowing by name. cwd sets the subprocess working directory and add_dirs grants access outside it. settings takes a path or inline JSON. plugins loads local bundles, each {"type": "local", "path": ...}. effort takes low, medium, high, xhigh or max, and is separate from extended thinking, set with thinking={"type": "adaptive" | "enabled" | "disabled"}; max_thinking_tokens is deprecated in favour of thinking. Similarly 'Skill' in allowed_tools is deprecated in favour of the skills option, which takes "all" or a list of names and validates it strictly, so skills=["*"] now raises before the CLI starts.
Here is the error table to keep beside you for your first week. Each row is a message you will genuinely see.
| What you see | Why | What to do |
|---|---|---|
error: externally-managed-environment |
pip against system Python | Use a venv or uv |
Claude Code not found at: /your/configured/path |
No bundled binary (sdist), wrong cli_path, or a different service PATH |
Install natively, fix cli_path, check claude --version in the service environment |
Native CLI binary for <platform>-<arch> not found |
Optional deps skipped, often npm ci --omit=optional |
Reinstall with optional deps or set pathToClaudeCodeExecutable |
Refusing to execute batch script 'C:\...\npm\claude.cmd' |
cli_path points at a .cmd/.bat shim |
Install natively and point at claude.exe |
Not logged in · Please run /login / Invalid API key |
Key not in the subprocess environment; the SDK does not read .env; in TS, env replaced process.env |
Export the key, load dotenv, spread process.env |
Not connected. Call connect() first. |
A client method used outside the lifecycle | Use async with ClaudeSDKClient() as client: |
Reached maximum number of turns |
Hit max_turns; raised after the error_max_turns result |
Catch it and resume with a higher limit |
Prompt is too long |
Context window exceeded | Compact, delegate to subagents, trim tool output |
| MCP tools visible but never called | Not pre-approved, so the flow denies them | Add mcp__server__tool or mcp__server__* to allowed_tools |
MCP server status: "failed" or "needs-auth" in init |
Missing env or tokens, package not installed, 30 s MCP_TIMEOUT, network |
Fix config, raise MCP_TIMEOUT, call reconnect_mcp_server() |
Skill <name> is not in this session's skills allowlist |
Skill not listed in skills |
Add it, or dispatch it with /<name> |
Command failed with exit code 1 with Error output: Check stderr output for details |
The CLI exited with no error result. That "Error output" text is fixed boilerplate, not real stderr | Pass a stderr= callback and log what it gives you |
API Error: Request rejected (429) |
Rate limit on the key or project | Lower concurrency, fan out fewer subagents, raise the tier |
API Error: Repeated 529 Overloaded errors |
API capacity; does not count against your quota | Retry later, or set fallback_model |
Credit balance is too low |
Console credits exhausted | Add credits and set workspace caps |
SELF_SIGNED_CERT_IN_CHAIN or UNABLE_TO_GET_ISSUER_CERT_LOCALLY |
A TLS-inspecting proxy sits in the path | Set NODE_EXTRA_CA_CERTS to the CA bundle. Not retried since CLI v2.1.199 |
Three of those deserve emphasis. The Error output: Check stderr output for details text is fixed boilerplate — people hunt for "the real stderr" for an hour before registering a callback. The TLS row matters wherever a managed inspection proxy is standard, which in practice means most Gulf enterprise networks; the cure is to point the runtime at the proxy's CA, never to disable verification. And status: "pending" is absent from the table because it is not an error.
- In TypeScript, set
env: { FOO: "bar" }without spreading and read the resulting error. - Run with
setting_sources=[]and compare theinittool list against the default. - Register a
stderrcallback and log everything it emits during a failing run.
Putting it all together: a reviewer that comments on a diff
One small end-to-end project that uses nearly everything above. The agent reads the working tree's diff, reviews it against the repository's own conventions, and returns structured findings. It may read and search, it may run exactly two git commands, it may not write, and it has a budget.
import asyncio
import json
from claude_agent_sdk import (
query,
ClaudeAgentOptions,
PermissionResultAllow,
PermissionResultDeny,
AssistantMessage,
ToolUseBlock,
ResultMessage,
ResultError,
)
SCHEMA = {
"type": "object",
"properties": {
"findings": {
"type": "array",
"items": {
"type": "object",
"properties": {
"file": {"type": "string"},
"severity": {"type": "string", "enum": ["low", "medium", "high"]},
"note": {"type": "string"},
},
"required": ["file", "severity", "note"],
},
},
"verdict": {"type": "string", "enum": ["approve", "comment", "request-changes"]},
},
"required": ["findings", "verdict"],
}
ALLOWED_COMMANDS = ("git diff", "git status")
async def gate(tool_name, input_data, context):
if tool_name == "Bash":
command = input_data.get("command", "").strip()
if not command.startswith(ALLOWED_COMMANDS):
return PermissionResultDeny(
message=f"Only {ALLOWED_COMMANDS} are permitted. Use Read or Grep instead."
)
return PermissionResultAllow()
PROMPT = """Review the uncommitted changes in this repository.
Read CLAUDE.md first if it exists, and judge the diff against the conventions it states.
Report only findings you can point at a specific file for."""
async def main():
options = ClaudeAgentOptions(
cwd=".",
system_prompt={"type": "preset", "preset": "claude_code"},
tools=["Read", "Glob", "Grep", "Bash"],
allowed_tools=["Read", "Glob", "Grep"],
disallowed_tools=["Write", "Edit", "WebFetch"],
permission_mode="default",
can_use_tool=gate,
max_turns=25,
max_budget_usd=0.50,
output_format={"type": "json_schema", "schema": SCHEMA},
)
try:
async for message in query(prompt=PROMPT, options=options):
if isinstance(message, AssistantMessage):
for block in message.content:
if isinstance(block, ToolUseBlock):
print(f" -> {block.name}")
elif isinstance(message, ResultMessage):
if message.terminal_reason and message.terminal_reason != "success":
raise SystemExit(f"terminal_reason={message.terminal_reason}")
if message.subtype != "success" or message.structured_output is None:
raise SystemExit(f"review failed: {message.subtype}")
print(json.dumps(message.structured_output, indent=2))
print(f"session={message.session_id} cost=${message.total_cost_usd or 0:.4f}")
except ResultError as exc:
raise SystemExit(f"agent error result: {exc}")
asyncio.run(main())
Walk through the choices, because each is a decision you will make again. system_prompt restores Claude Code's preset, because reviewing a codebase is the task that prompt was written for. tools includes Bash so git is reachable, but allowed_tools deliberately omits it, so every bash call falls through to gate — the allowed_tools-is-not-a-restriction lesson used on purpose. disallowed_tools removes Write, Edit and WebFetch entirely, so a reviewer cannot modify the code it reviews. max_turns and max_budget_usd bound the run in both dimensions, and output_format gives a CI job something to act on. The result handler checks terminal_reason, then subtype, then structured_output, in that order, and the ResultError clause catches the raise that follows an error result in single-shot mode.
Run it in a repository with uncommitted changes and you get a JSON verdict, a resumable session ID, and a cost figure. From here the extensions are short: post the findings to a pull request, or wrap the whole thing in a container a CI pipeline can call — the Docker guide covers the packaging. Export OpenTelemetry and every run becomes a trace, which is where Langfuse or any OTLP backend fits; the CLI exports telemetry itself and the SDK passes the environment variables through, with one hard rule — never use the console exporter, because stdout is the SDK's message channel.
- Run
review.pyon a branch with real changes and read the findings. - Ask it to run
git pushin the prompt and watchgaterefuse. - Resume the session with a follow-up question about the highest-severity finding.
What you can now do, and what comes next
You can install the SDK on macOS, Linux or Windows and diagnose the failures specific to each. You can run an agent that reads and modifies real files, read the message stream, and tell a genuine success from an apparent one by checking terminal_reason before subtype. You can shape the tool surface with tools, disallowed_tools and allowed_tools, and explain why the third is not a restriction. You can set a permission mode, write an approval callback, and recognise when it is shadowed. You can continue, resume and fork sessions, and you know a session restores a conversation and not a filesystem. You can keep a client alive with ClaudeSDKClient and interrupt it, write a custom tool and name it correctly, and demand JSON and validate that you got it.
Three habits will serve you from here. Pin the SDK version, because that also pins the bundled CLI, then take patches continuously and read the changelog before any minor — the 0.2 to 0.3 TypeScript bump removed the entire V2 session API. Log session_id on every run. And set a budget on every agent, because both caps are off by default and the failure mode of an unbounded agent is financial.
What the Mid-level guide takes up: the agent loop in enough detail to predict its behaviour, subagents and the Agent tool, hooks across the full event list, external MCP servers and tool search, session storage adapters that mirror transcripts so another host can resume, the hosting patterns, and real cost control. The Senior guide covers sandboxing and the security model, multi-tenancy in a shared container, egress proxies that inject the API key so the agent never sees it, and where the SDK stops being the right tool.
Useful neighbours in this catalogue: Claude API for the layer underneath, MCP for the tool protocol, OpenAI Agents SDK and CrewAI for how other frameworks divide the same problem, Langfuse and OpenTelemetry for observability, AWS Bedrock for regional model hosting, and Docker for shipping any of it.
Sources
- Agent SDK overview
- Quickstart
- Migration guide
- Troubleshooting
- Configuration
- The agent loop
- Sessions
- Streaming vs single mode
- Structured outputs
- Custom tools
- MCP in the Agent SDK
- Permissions
- Hooks
- Modifying system prompts
- Cost tracking
- Observability
- TypeScript reference
- Python reference
- Errors
- Environment variables
- TypeScript SDK changelog
- Python SDK changelog