This is part one of three. It covers everything you need to build a real LangGraph application, not a demo you throw away. By the end you can define a graph's state, write nodes, branch between them, give a conversation memory that survives a restart, stream output token by token, pause a run to ask a human for approval, and read the error messages LangGraph produces when you get it wrong. Mid-level and Senior take the same topics further; nothing here is wasted.
Everything below was checked against LangGraph 1.2.12, released 21 September 2026. LangGraph follows semantic versioning, so anything on the 1.2 line behaves the same way. Where the official documentation and the shipped code disagree, this guide says so rather than picking one quietly.
Each section ends with a Try it task. Do them as you go. Graphs are one of those topics that feel abstract on the page and obvious the moment you have watched your own graph loop, pause, and resume.
What LangGraph is, and the problem it solves
LangGraph is a low-level orchestration framework and runtime for building long-running, stateful agents. You describe your application as a graph: a shared piece of data called the state, a set of functions called nodes that each read the state and return an update to it, and edges that decide which node runs next. LangGraph runs that graph, saves the state after every step, and lets you stop, inspect, rewind and resume it.
Notice what is not in that description. LangGraph does not write your prompts, does not pick your model, and does not decide what an agent is. It has no opinion about whether your application is an agent at all. It is the machinery underneath: the part that remembers where you were and decides what happens next. LangChain's create_agent is the higher-level agent framework built on top of it, and LangSmith is the separate product that traces, evaluates and deploys what you build. You can use LangGraph with none of those.
The problem it solves becomes obvious the second time you write an LLM application. The first one is a function: take a question, call a model, return the answer. The second one needs to call a tool, look at the result, maybe call another tool, maybe ask the user a clarifying question, and stop when it is finished. Written by hand, that becomes a while True: loop around a growing list of messages, with a pile of if statements inside it.
That loop works until it meets reality, and then four things go wrong at once.
It forgets everything when the process ends. The conversation lives in a local variable. Restart your server, deploy a new version, or let a request time out, and the user starts over. For a chatbot that is annoying. For a workflow that has already charged a customer's card, it is a bug with a cost attached.
You cannot pause it. Any workflow that needs a human to approve something — a refund, an outbound email, a database migration — has to stop mid-flight, wait an unknown length of time, and then carry on from exactly where it was. A while loop in a request handler cannot wait three hours.
You cannot see inside it. When the loop calls the wrong tool on step seven, you have a list of messages and a guess. There is no record of what the state looked like at step six.
Every retry starts from zero. If the fourth of five model calls fails, you pay for the first three again.
LangGraph's answer is to make the loop a data structure instead of control flow. Because the graph is described rather than executed inline, the runtime can checkpoint it, pause it, replay it, visualise it, and stream from it.
The lineage is worth knowing, because it explains the vocabulary. LangGraph's execution model comes from Pregel, Google's system for large-scale graph processing, and from Apache Beam. That is where the word "super-step" comes from, and it is why nodes communicate by writing to shared channels rather than by calling each other. The public API deliberately resembles NetworkX: you add nodes, you add edges, you compile.
- Take an LLM script you have written, or imagine one that answers questions using a search tool.
- Write down every point at which it makes a decision about what to do next.
- For each one, write down what information that decision depends on.
Graphs versus the loop you would have written
Before committing to the graph model, it is worth being honest about what you give up, because LangGraph is not free. A graph is more code than a loop for a small task, and the indirection is real: you cannot read a graph top to bottom the way you read a function.
Hand-written loop
- Reads top to bottom, no framework to learn
- State lives in local variables, lost on exit
- Pausing for a human means inventing your own storage
- Retrying means rerunning everything
- Parallel steps are your problem (threads, asyncio, merging)
- Debugging means print statements
LangGraph
- More structure up front: a state schema, nodes, edges
- State is saved after every step, keyed by a thread
interrupt()pauses a run and returns a payload- Resume continues from the last checkpoint
- Nodes triggered together run in parallel and merge through reducers
- Every past state is inspectable and replayable
The honest rule for choosing: if your application is one model call with no branching and no memory, write the function. The moment it has a loop, needs to remember a conversation, or has to survive a restart, the graph pays for itself. Most teams arrive at LangGraph from the second direction, having already written the loop and discovered its limits.
There is a second axis you will meet in the documentation and should know exists even though this guide does not teach it. LangGraph offers two front ends over the same runtime. The Graph API is what you are about to learn: StateGraph, nodes and edges, written out explicitly. The Functional API uses @entrypoint and @task decorators with ordinary Python control flow — if, for, while — and gets you checkpointing and interrupts without drawing a graph. It compiles down to a single node, which has a consequence you need to know before you reach for it: on resume it replays your function from the beginning, using cached results for tasks that already completed. Anything outside a task must therefore be deterministic. Start with the Graph API. It makes the execution model visible, and the visible version is the one you can debug.
langgraph-cli, still langgraph.json.
- For the decision list you wrote in the last section, count the decisions.
- Ask: does anything here need to outlive a single HTTP request?
- Write one sentence saying whether a graph is worth it, and why.
The four nouns: state, node, edge, super-step
Almost everything in LangGraph is built from four ideas. Learn them precisely now and the rest of the API will look like consequences rather than new material.
State is the shared data structure that is passed between nodes. You declare its shape with a schema, and the most common choice is a TypedDict: a dictionary whose keys and value types are written down. You can also use a dataclass, which gives you default values, or a Pydantic BaseModel, which validates values at runtime at some cost in speed. The state is not a bag of globals — it is the single place your graph keeps what it knows.
A node is a Python function, synchronous or asynchronous, that takes the state and returns a partial update: a dictionary containing only the keys it wants to change. This is the single most important detail in the whole model. A node does not return the new state. It returns a small dictionary of changes, and LangGraph merges them. A node that only sets a counter returns {"count": 3} and says nothing about the other twelve keys.
An edge says what runs next. A normal edge, add_edge("a", "b"), always goes from a to b. A conditional edge, add_conditional_edges("a", router), calls your router function with the state; the router returns the name of the next node, a list of names, or END. Two virtual nodes mark the boundaries: START is where input enters, and END is where the run finishes. They are not functions you write, just names.
A super-step is one tick of the runtime. Every node triggered in the same tick runs in parallel, and their updates are applied together at the end of the step. One checkpoint is written per super-step. This explains behaviour that is otherwise baffling: if two parallel nodes both write the same state key in one step, LangGraph does not know which should win, and it refuses rather than guessing.
node A
nodes B and C in parallel
Two more words you will see immediately and should not be scared of. A channel is the low-level storage slot behind one state key; the default channel type simply overwrites, which is why an ordinary key behaves like a variable. A reducer is a function you attach to a key to say how updates combine instead of overwriting. The next-but-one section is entirely about reducers, because they cause most beginner confusion.
It is worth pausing on why the super-step exists at all, because it is the one idea with no equivalent in ordinary Python. In a normal program, a function's effects are visible the instant it returns. In a Pregel-style runtime, effects are collected and applied at a boundary. That buys two things. It makes parallelism safe without locks, because no node ever observes another node's half-finished write. And it gives the runtime a natural place to save: one consistent snapshot per step, rather than a smear of partial updates. Every feature in the rest of this guide — memory, resume, replay, streaming — is built on that boundary existing.
The practical consequence for your code is a small discipline. A node should read the state it was given and return its update. It should not expect to see a sibling node's write within the same step, and it should not reach outside the state to communicate. When you find yourself wanting a node to see something another node just produced, the answer is almost always to put the second node in a later step with an edge, not to work around the boundary.
Finally, compile. You build a graph with StateGraph, then call .compile(). Compiling validates the structure — it catches orphan nodes, missing entry points and edges pointing at names that do not exist — and returns a CompiledStateGraph. That object is a LangChain Runnable, which is where invoke, stream, batch and their async twins come from. You must compile before you can run anything.
- Say out loud what a node returns. If the answer was "the new state", say it again correctly.
- Sketch a three-node graph on paper with one branch in it.
- Mark which super-step each node runs in.
Installing LangGraph and checking the setup
LangGraph is a pure-Python package with ordinary dependencies, so installation is unremarkable on Linux, macOS and Windows. The one requirement to note is Python: the library needs Python 3.10 or newer, and the command-line tool's local development server needs 3.11 or newer, because it pulls in the Agent Server package.
Start with a virtual environment. On Linux or macOS:
python3 -m venv .venv
source .venv/bin/activate
pip install -U langgraph
On Windows with PowerShell:
py -3.12 -m venv .venv
.venv\Scripts\Activate.ps1
pip install -U langgraph
If PowerShell refuses to run the activation script, Set-ExecutionPolicy -Scope CurrentUser RemoteSigned once will fix it. If you prefer uv, uv venv && source .venv/bin/activate && uv add langgraph does the same job faster.
That one package is enough to build and run graphs. Models and tools are separate, because LangGraph does not depend on any model provider:
pip install -U langchain # create_agent, model wrappers, tool helpers
pip install -U langchain-openai # or langchain-anthropic, etc.
For a conversation that survives a restart you want a real checkpointer. SQLite is the right choice while you are learning, and Postgres is the one you will use at work:
pip install -U langgraph-checkpoint-sqlite
pip install -U "psycopg[binary,pool]" langgraph-checkpoint-postgres
Finally, the CLI and the local development server. The quotes around the extras matter — zsh, which is the default shell on macOS, treats square brackets as a glob pattern and will tell you there is no match:
pip install -U "langgraph-cli[inmem]"
Now verify all of it. Three checks, each of which fails loudly if something is wrong:
python -c "import importlib.metadata as m; print(m.version('langgraph'))"
# 1.2.12
langgraph --help
# usage: langgraph [OPTIONS] COMMAND [ARGS]...
langgraph dev
# starts on http://127.0.0.1:2024, prints a Studio link, serves OpenAPI at /docs
curl http://127.0.0.1:2024/ok
# {"ok":true}
langgraph dev runs the server in process with in-memory storage and does not need Docker. The commands that build container images — langgraph build and langgraph up — do. If langgraph dev opens a browser you cannot use, Safari and locked-down corporate networks are the usual culprits; langgraph dev --tunnel serves it through a public tunnel instead. Its other useful flags are --port (default 2024), --host (default 127.0.0.1), --no-browser and --no-reload.
pip and uv can fail with certificate errors that have nothing to do with LangGraph. With uv, --native-tls makes it use the operating system's trust store, which already contains your employer's certificate. This bites people in Gulf and Egyptian enterprises with managed devices more often than it bites people on personal machines.
- Create a virtual environment and install
langgraphand"langgraph-cli[inmem]". - Run all three verification commands above.
- Run
langgraph devand open the printed Studio link.
{"ok":true}. Leave it running; you will point it at a real graph shortly.
Your first graph, step by step
We will build the smallest useful graph: one node that produces a message. No model key is needed, because a fake model teaches the mechanics better than a real one — when something goes wrong you know it is your graph and not the network.
from langgraph.graph import StateGraph, MessagesState, START, END
def mock_llm(state: MessagesState):
return {"messages": [{"role": "ai", "content": "hello world"}]}
builder = StateGraph(MessagesState)
builder.add_node(mock_llm)
builder.add_edge(START, "mock_llm")
builder.add_edge("mock_llm", END)
graph = builder.compile()
result = graph.invoke({"messages": [{"role": "user", "content": "hi!"}]})
for message in result["messages"]:
message.pretty_print()
Running it prints the user message followed by the AI message. Six lines of graph code, and every one of them is doing something worth naming.
StateGraph(MessagesState) creates a builder whose state schema is MessagesState, a prebuilt schema with a single key, messages, that holds a list of messages and knows how to append to it. You will use it constantly.
builder.add_node(mock_llm) registers the function as a node. Passing the function alone takes the node's name from the function's name — hence "mock_llm" in the edges below. You can be explicit instead: builder.add_node("call_model", mock_llm). Be deliberate, because the name is what edges, streamed updates, checkpoints and traces all refer to.
The two add_edge calls connect the virtual START node to your node and your node to END. Without the first one, compiling fails: a graph with no path from START has no way to begin.
builder.compile() validates and returns the runnable graph. graph.invoke(...) runs it to completion and returns the final state.
Look closely at what the node returned: {"messages": [...]}, a list with one new message — not the whole conversation. The user's message is still in the result because MessagesState's messages key appends rather than overwrites. That append is a reducer, and the next section is about what happens when you forget to ask for one.
Now make it a graph with more than one step, and with state of your own design:
from typing import TypedDict
from langgraph.graph import StateGraph, START, END
class State(TypedDict):
topic: str
draft: str
word_count: int
def write_draft(state: State) -> dict:
return {"draft": f"A short note about {state['topic']}."}
def count_words(state: State) -> dict:
return {"word_count": len(state["draft"].split())}
builder = StateGraph(State)
builder.add_node("write_draft", write_draft)
builder.add_node("count_words", count_words)
builder.add_edge(START, "write_draft")
builder.add_edge("write_draft", "count_words")
builder.add_edge("count_words", END)
graph = builder.compile()
print(graph.invoke({"topic": "checkpointers"}))
# {'topic': 'checkpointers', 'draft': 'A short note about checkpointers.', 'word_count': 5}
Two super-steps, two checkpoints' worth of history if a checkpointer were attached, and a final state containing everything. Note that count_words reads state["draft"], a key it never wrote. That is the whole point of shared state: nodes communicate through it rather than by calling each other.
One habit to adopt immediately: have the graph draw itself.
print(graph.get_graph().draw_mermaid())
That prints Mermaid source you can paste into any Markdown viewer that renders it. draw_mermaid_png() gives you an image instead. For a graph with branches, looking at the picture catches wiring mistakes in seconds that reading code does not.
- Run
first_graph.pyand thentwo_nodes.py. - Add a third node to the second graph that uppercases the draft, and wire it between the two existing nodes.
- Print the Mermaid diagram and check the order is what you intended.
Reducers, and why your list keeps being overwritten
Here is the bug every beginner writes. A graph collects findings from several nodes, and only the last one survives.
class State(TypedDict):
findings: list[str] # no reducer: this key is OVERWRITTEN
def search_docs(state: State) -> dict:
return {"findings": ["found in the docs"]}
def search_code(state: State) -> dict:
return {"findings": ["found in the code"]}
Run these one after the other and findings ends up as a single-item list, because the default channel behind a state key simply stores the last value written. Run them in parallel in the same super-step and you get something more useful: an error.
langgraph.errors.InvalidUpdateError: At key 'findings': Can receive only one value per step.
Use an Annotated key to handle multiple values.
For troubleshooting, visit: https://docs.langchain.com/oss/python/langgraph/errors/INVALID_CONCURRENT_GRAPH_UPDATE
That message is telling you exactly what to do. A reducer is a function (current, update) -> new that you attach to a key with Annotated, and it tells LangGraph how to combine writes instead of replacing them:
import operator
from typing import Annotated, TypedDict
class State(TypedDict):
findings: Annotated[list[str], operator.add]
operator.add on two lists concatenates them, so now both nodes' findings are kept, in parallel or in sequence. The reducer runs at the super-step boundary, where all the step's writes are applied together.
For messages there is a purpose-built reducer, add_messages, and MessagesState is simply a state that uses it:
from typing import Annotated, TypedDict
from langchain_core.messages import AnyMessage
from langgraph.graph.message import add_messages
class State(TypedDict):
messages: Annotated[list[AnyMessage], add_messages]
add_messages does three things that plain concatenation cannot. It appends new messages. It replaces by message ID, so re-emitting a message with the same ID edits it rather than duplicating it. And it accepts plain dictionaries such as {"role": "user", "content": "hi"} as well as message objects, converting them for you. That last point is why the first example in this guide could pass a dictionary.
Now the trap that follows from all of this, and it catches experienced people too.
return {"findings": []} runs the reducer, and concatenating an empty list changes nothing. To replace rather than merge, bypass the reducer explicitly with Overwrite:
from langgraph.types import Overwrite
def reset(state: State) -> dict:
return {"findings": Overwrite([])}
Read it as a rule you can apply without thinking: a key with no reducer behaves like a variable, a key with a reducer behaves like an accumulator, and Overwrite is how you assign to an accumulator.
- Build a graph with two nodes that both write
findings, fanned out in parallel fromSTART. - Run it without a reducer and read the full
InvalidUpdateError, including the troubleshooting URL. - Add
Annotated[list[str], operator.add]and run it again. - Add a node that tries to clear the list with
[], see it fail, and fix it withOverwrite.
Branching: conditional edges and loops
Straight lines are rarely what you want. A conditional edge gives the graph a decision: a function that reads the state and returns where to go next.
from typing import Annotated, TypedDict
import operator
from langgraph.graph import StateGraph, START, END
class State(TypedDict):
question: str
attempts: Annotated[int, operator.add]
answer: str
def attempt(state: State) -> dict:
return {"attempts": 1, "answer": "not sure yet"}
def finalise(state: State) -> dict:
return {"answer": f"answered after {state['attempts']} attempts"}
def route(state: State) -> str:
if state["attempts"] >= 3:
return "finalise"
return "attempt"
builder = StateGraph(State)
builder.add_node("attempt", attempt)
builder.add_node("finalise", finalise)
builder.add_edge(START, "attempt")
builder.add_conditional_edges("attempt", route, ["attempt", "finalise"])
builder.add_edge("finalise", END)
graph = builder.compile()
print(graph.invoke({"question": "why?", "attempts": 0, "answer": ""}, {"recursion_limit": 10}))
# {'question': 'why?', 'attempts': 3, 'answer': 'answered after 3 attempts'}
Three details carry the weight here.
The router is a plain function. It receives the state and returns a node name, a list of node names, or END. It must not return anything else; a typo in the returned string becomes a runtime error rather than a wrong branch.
The third argument to add_conditional_edges lists the possible destinations. You can pass a list of node names, or a dictionary mapping the router's return values to node names when the two differ. It is optional for execution but not for drawing: without it, the rendered diagram cannot show where the branch can go.
The graph loops. attempt routes back to attempt, which is perfectly legal and is how every agent works. The exit condition lives in the router, and if you get it wrong the graph runs until the recursion limit stops it.
GraphRecursionError: "Recursion limit of 5 reached without hitting a stop condition. You can increase the limit by setting the recursion_limit config key." The defaults in 1.x are deliberately high — the docs say 1000 from version 1.0.6, while the shipped 1.2.12 code uses 10007 and reads the environment variable LANGGRAPH_DEFAULT_RECURSION_LIMIT. Either way, a runaway loop will burn thousands of model calls before it stops. Pass {"recursion_limit": 10} or similar on every graph that cycles. If you learned that the default was 25, that was true before 1.0.6 and is not true now.
There is a second way to route, and you should recognise it even though you will not need it on day one. Returning a Command object from a node combines the state update and the routing decision in one place:
from typing import Literal
from langgraph.types import Command
def attempt(state: State) -> Command[Literal["attempt", "finalise"]]:
if state["attempts"] >= 2:
return Command(update={"attempts": 1}, goto="finalise")
return Command(update={"attempts": 1}, goto="attempt")
Use conditional edges when the decision is about the graph's shape and belongs outside the node. Use Command when the node itself has just learned something that decides where to go, typically because a model told it. Both are first-class; neither is deprecated.
- Run the router graph, then delete the
attempts >= 3check so it never exits. - Run it with
{"recursion_limit": 5}and read theGraphRecursionError. - Fix the router, then rewrite the same logic using
Commandand compare which version you find clearer.
Memory: checkpointers and threads
So far every invoke has started from nothing. To give a graph memory you attach a checkpointer at compile time and pass a thread ID at call time. Those two things together are LangGraph's entire short-term memory story.
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import StateGraph, MessagesState, START, END
def echo(state: MessagesState) -> dict:
last = state["messages"][-1].content
return {"messages": [{"role": "ai", "content": f"you said: {last}"}]}
builder = StateGraph(MessagesState)
builder.add_node(echo)
builder.add_edge(START, "echo")
builder.add_edge("echo", END)
graph = builder.compile(checkpointer=InMemorySaver())
config = {"configurable": {"thread_id": "conversation-1"}}
graph.invoke({"messages": [{"role": "user", "content": "first"}]}, config)
result = graph.invoke({"messages": [{"role": "user", "content": "second"}]}, config)
print(len(result["messages"]))
# 4
Four messages, not two. The second invoke loaded the saved state for conversation-1, appended to it, and saved it again. Change the thread ID and you get a fresh conversation with no shared history. That is how one deployed graph serves thousands of independent users: one thread each.
A checkpoint is a snapshot of the state at the end of a super-step. A thread is the ordered sequence of checkpoints under one thread_id. The checkpointer is the thing that stores them.
Forget the thread ID and the error is specific:
ValueError: Checkpointer requires one or more of the following 'configurable' keys:
thread_id, checkpoint_ns, checkpoint_id
InMemorySaver (also exported as MemorySaver) keeps everything in RAM, which makes it perfect for tests and useless for anything else — restart the process and every conversation is gone. For local work that should survive, use SQLite. For production, use Postgres:
from langgraph.checkpoint.postgres import PostgresSaver
DB_URI = "postgresql://user:pass@localhost:5432/langgraph"
with PostgresSaver.from_conn_string(DB_URI) as checkpointer:
checkpointer.setup() # required the first time: creates the tables
graph = builder.compile(checkpointer=checkpointer)
graph.invoke({"messages": [...]}, config)
setup() is the step people forget, and forgetting it produces missing-table errors from Postgres rather than a helpful LangGraph message. Call it once, on first use, for the checkpointer and for the store. If you are writing async code, use AsyncPostgresSaver with ainvoke and astream; mixing a synchronous saver into an async graph is a reliable way to confuse yourself.
Because every step is saved, you can look at it:
snapshot = graph.get_state(config)
print(snapshot.values) # the state right now
print(snapshot.next) # which nodes run next; () means finished
for past in graph.get_state_history(config):
print(past.metadata["step"], past.values) # newest first
A StateSnapshot carries the values, what runs next, the config that identifies it (thread, namespace, checkpoint ID), metadata including the step number and the writes that produced it, and the pending tasks. This is the thing that replaces print-statement debugging. When an agent misbehaved on step seven, you can read step six.
One knob you should know the name of without tuning it yet. Durability controls when checkpoints are actually written: "exit" saves only when the run finishes and is fastest, "async" writes while the next step runs and is the default, and "sync" writes before the next step starts and is safest. Pass it per call as graph.invoke(x, config, durability="sync"). If you see checkpoint_during=True in an older example, that parameter is deprecated and durability replaced it; passing both raises an error.
InMemoryStore to start with, PostgresStore in production) with put, get, search and delete. Mid-level covers it properly. For now just do not try to make the checkpointer do that job.
- Run the memory graph twice on the same thread ID and count the messages.
- Run it on a new thread ID and count again.
- Drop the
configargument entirely and read theValueError. - Print
get_state_history(config)and identify the step number of each checkpoint.
Streaming: watching the graph work
invoke waits for the whole graph and hands back the final state. That is fine for a batch job and unacceptable for a user interface, where a blank screen for twenty seconds reads as a broken application. stream yields as the graph runs, and stream_mode chooses what you get.
for chunk in graph.stream({"messages": [{"role": "user", "content": "hi"}]}, config, stream_mode="updates"):
print(chunk)
# {'echo': {'messages': [AIMessage(content='you said: hi')]}}
The modes worth knowing on day one:
stream_mode |
What each item is | Use it for |
|---|---|---|
"values" |
the full state after each super-step | debugging, progress snapshots |
"updates" |
only what each node returned, keyed by node name | logging which node did what |
"messages" |
LLM tokens as (chunk, metadata) pairs |
typing-effect chat output |
"custom" |
whatever you write yourself | progress for long non-LLM work |
"debug" |
detailed execution events | hard problems |
There are also "checkpoints" and "tasks" for inspecting persistence and scheduling. You can pass a list — stream_mode=["values", "messages"] — and each yielded item then tells you which mode it came from.
The "custom" mode is the one beginners underuse. A node that spends thirty seconds in a non-LLM operation can report progress itself:
from langgraph.config import get_stream_writer
def index_documents(state: State) -> dict:
writer = get_stream_writer()
for i, doc in enumerate(state["docs"]):
writer({"indexed": i + 1, "total": len(state["docs"])})
return {"indexed": len(state["docs"])}
get_stream_writer() does not work inside async nodes. Context does not propagate into asyncio tasks before 3.11, so take a writer parameter in the node signature instead, and pass config explicitly into async model calls such as model.ainvoke(..., config). On 3.11 and newer this is not a concern. It is one more reason to start on 3.12.
A question that comes up immediately: which mode should a chat interface use? In practice, two at once. "messages" drives the typing effect, because it yields model tokens as they arrive, and "updates" tells your interface which node is running so you can show something honest like "searching" or "writing the answer". Streaming only "values" to a browser is tempting because it is simple, but you are then sending the entire conversation state on every step, which gets expensive as a thread grows. Stream the smallest thing that answers the question your interface is asking.
Two things about output shapes will save you confusion when reading the documentation. First, invoke and stream take a version argument, and the default is still "v1". Version "v2", added in 1.1, returns a typed GraphOutput object with .value and .interrupts instead of a plain dictionary. It is opt-in, so everything in this guide uses the v1 shapes. Second, there is a newer, separate method — stream_events(..., version="v3") — with convenient projections such as run.messages and run.values. Current documentation pages often show it and call it recommended, but it is still labelled beta. Learn stream first; it is what the stable API gives you.
- Run your memory graph with
stream_mode="updates", then with"values", and describe the difference in one sentence. - Pass
stream_mode=["values", "updates"]and look at what arrives. - Add a node that emits custom progress with
get_stream_writer()and stream with"custom".
Pausing for a human: interrupt
This is the feature that makes LangGraph worth learning even for workflows that barely involve a model. interrupt() stops the graph in the middle of a node, saves the state, and hands a payload back to the caller. Later — seconds or days later — you resume with a value, and that value becomes the return value of the interrupt() call.
from typing import TypedDict
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import StateGraph, START, END
from langgraph.types import Command, interrupt
class State(TypedDict):
amount: int
decision: str
def request_approval(state: State) -> dict:
answer = interrupt({"question": f"Approve a refund of {state['amount']}?"})
return {"decision": answer}
builder = StateGraph(State)
builder.add_node("request_approval", request_approval)
builder.add_edge(START, "request_approval")
builder.add_edge("request_approval", END)
graph = builder.compile(checkpointer=InMemorySaver())
config = {"configurable": {"thread_id": "refund-42"}}
first = graph.invoke({"amount": 500, "decision": ""}, config)
print(first["__interrupt__"])
# (Interrupt(value={'question': 'Approve a refund of 500?'}, ...),)
final = graph.invoke(Command(resume="approved"), config)
print(final["decision"])
# approved
The first invoke returns early, with the payload under the __interrupt__ key of the result (that is the v1 output shape; with version="v2" you read result.interrupts instead). The process can now exit entirely. The second invoke passes Command(resume="approved") on the same thread ID, the graph loads its checkpoint, and interrupt() returns "approved".
The rules around interrupt are short, strict, and all follow from one fact: the node re-runs from the beginning on resume. interrupt() works by raising an internal exception, and resuming replays the node with the answer cached.
- Never wrap
interrupt()intry/except. A bareexcept Exception:swallows the pause and the graph carries on with nonsense. - Make side effects before an
interrupt()idempotent, or move them after it. Code above the call runs again on resume, so an email sent there is sent twice. - Do not reorder
interrupt()calls inside a node between runs. Resume values are matched by position. - Keep payloads simple and JSON-serialisable. They cross a process boundary.
And the trap that wastes an afternoon: interrupt() needs a checkpointer and a thread ID, but calling it without a checkpointer does not raise an error. The graph returns {'__interrupt__': [...]} exactly as it would otherwise, and there is simply nothing to resume. If a resume appears to do nothing, check that you passed checkpointer= to compile().
New in 1.2.12, you can have the resume value validated instead of trusting whatever arrives, by passing a Pydantic model, TypedDict or dataclass as response_schema:
from pydantic import BaseModel
from langgraph.types import interrupt
class Decision(BaseModel):
approved: bool
note: str = ""
def request_approval(state) -> dict:
decision = interrupt({"question": "Approve?"}, response_schema=Decision)
return {"decision": "yes" if decision.approved else "no"}
Related but different: interrupt_before=["node"] and interrupt_after=["node"], which you can pass at compile or call time. These are static breakpoints — they pause between nodes rather than inside one, and they exist for debugging and for stepping through a graph in Studio, not for building approval flows.
- Run
approval.pyand read theInterruptobject it returns. - Add
print("charging the card")above theinterrupt()call and run both halves again. Count how many times it prints. - Remove
checkpointer=and watch the resume silently fail.
An agent without writing the loop
Everything so far was hand-built, which is the right way to learn. But the single most common thing people build — a model that can call tools in a loop until it is finished — already exists, and writing it yourself every time is a waste.
from langchain.agents import create_agent
def get_weather(city: str) -> str:
"""Return the weather for a city."""
return f"It is 34 degrees and sunny in {city}."
agent = create_agent(
"openai:gpt-4.1",
tools=[get_weather],
system_prompt="You are a concise assistant. Use the tools you have.",
)
result = agent.invoke({"messages": [{"role": "user", "content": "weather in Cairo?"}]})
print(result["messages"][-1].content)
create_agent returns a compiled LangGraph graph. Everything you have learned still applies: pass a checkpointer for memory, pass thread_id in the config, stream it, inspect it with get_state. It is not a separate world, it is a graph someone wrote for you.
from langchain.agents import create_agent, with system_prompt=. The older from langgraph.prebuilt import create_react_agent, with prompt=, was deprecated in LangGraph 1.0. It still imports in 1.2.12, which is why so much tutorial code uses it, but new code should not. Several neighbours were deprecated at the same time: MessageGraph (use StateGraph with a messages key), the NodeInterrupt exception (use interrupt()), ValidationNode, and the agent state classes in langgraph.prebuilt. Also prefer from langgraph.types import Send, Interrupt over the old langgraph.constants path; START and END are fine from either.
So when do you build the graph yourself? When you need control the agent loop does not give you: a fixed sequence of steps, a human approval in the middle, a branch that depends on business rules rather than the model's judgement, or a shape you want to be able to draw for a compliance reviewer. A good instinct is to reach for create_agent when the model should decide what happens next, and for StateGraph when you should decide.
- Run the agent with a real provider key, and watch the tool get called.
- Recompile it with a checkpointer and a thread ID, then ask a follow-up that depends on the first answer.
- Print
agent.get_graph().draw_mermaid()and find the loop between the model and the tools.
Reading LangGraph's error messages
LangGraph's errors are unusually good: most carry an error code and a link to a page about that exact failure. Learning to read them is faster than learning to avoid them. Here are the ones you will actually hit, with the real text.
InvalidUpdateError: At key 'k': Can receive only one value per step. Use an Annotated key to handle multiple values. Two parallel nodes wrote the same key and it has no reducer. Add one, or stop writing the key from two places at once.
GraphRecursionError: Recursion limit of 5 reached without hitting a stop condition. A cycle with no exit, or a genuinely long graph. Fix the router first — most of the time it should have returned END — and only then consider raising recursion_limit.
InvalidUpdateError: Expected dict, got ['whoops'] A node returned something that is not a dictionary of state keys. Usually one branch of an if forgot to return, or returned a bare list. Every code path must return a dict or a Command.
ValueError: Graph must have an entrypoint: add at least one edge from START to another node You forgot add_edge(START, "first"). Compiling catches it, which is exactly why compiling exists.
ValueError: Found edge ending at unknown node 'nope' A typo, or an edge added with a name that does not match the node's registered name. This is the usual punishment for letting add_node(fn) choose names implicitly and then renaming the function.
ValueError: Checkpointer requires one or more of the following 'configurable' keys: thread_id, checkpoint_ns, checkpoint_id You compiled with a checkpointer and invoked without a thread ID.
ValueError: Node timeouts are only supported for async nodes because sync Python execution cannot be safely cancelled in-process. Node 'a' is sync. You put a timeout on a def node. Make it async def. Timeouts are a 1.2 feature covered at Senior level, but the message is worth recognising.
Blocking-I/O warnings under langgraph dev. The development server detects synchronous blocking calls inside async code — requests.get, a synchronous database driver — and complains, because one blocked node can stall the whole event loop. The right fix is async clients (httpx.AsyncClient, psycopg in async mode) or asyncio.to_thread. langgraph dev --allow-blocking silences it for local work only.
Deprecation warnings are worth reading rather than filtering. They are specific and they tell you the replacement: "MessageGraph is deprecated in LangGraph v1.0.0, to be removed in v2.0.0. Please use StateGraph with a 'messages' key instead.", "'config_schema' is deprecated and will be removed. Please use 'context_schema' instead.", "'checkpoint_during' is deprecated and will be removed. Please use 'durability' instead." Deprecated features keep working for at least one minor release and are removed only in a major version, so you have time — but a warning you fix today is not a migration you do under pressure later.
A debugging order that works for almost all of these. Read the error code and open the link. Print graph.get_state(config) to see where the graph actually stopped and what next says. Re-run with stream_mode="updates" and watch which node produced the bad value. If the shape of the graph is in doubt, print the Mermaid diagram.
- Deliberately cause three of the errors above and read each message to the end, including the URL.
- For each one, write the fix in your own words in a single line.
- Keep that file. It is the cheat sheet you will reach for in six weeks.
Putting it all together
One small project that uses everything: a support-ticket triage graph that classifies a ticket, drafts a reply, asks a human to approve anything it marks as a refund, and remembers the thread.
from typing import Annotated, TypedDict
import operator
from langgraph.checkpoint.sqlite import SqliteSaver
from langgraph.graph import StateGraph, START, END
from langgraph.types import Command, interrupt
class State(TypedDict):
ticket: str
category: str
draft: str
log: Annotated[list[str], operator.add]
sent: bool
def classify(state: State) -> dict:
text = state["ticket"].lower()
category = "refund" if "refund" in text or "money back" in text else "general"
return {"category": category, "log": [f"classified as {category}"]}
def draft_reply(state: State) -> dict:
draft = f"Thanks for getting in touch about your {state['category']} request."
return {"draft": draft, "log": ["drafted a reply"]}
def human_approval(state: State) -> dict:
answer = interrupt({"draft": state["draft"], "category": state["category"]})
return {"log": [f"human said {answer}"], "sent": answer == "approve"}
def send(state: State) -> dict:
# In real code this is the side effect, and it lives AFTER any interrupt.
return {"sent": True, "log": ["sent without approval"]}
def route(state: State) -> str:
return "human_approval" if state["category"] == "refund" else "send"
builder = StateGraph(State)
builder.add_node("classify", classify)
builder.add_node("draft_reply", draft_reply)
builder.add_node("human_approval", human_approval)
builder.add_node("send", send)
builder.add_edge(START, "classify")
builder.add_edge("classify", "draft_reply")
builder.add_conditional_edges("draft_reply", route, ["human_approval", "send"])
builder.add_edge("human_approval", END)
builder.add_edge("send", END)
with SqliteSaver.from_conn_string("triage.db") as checkpointer:
graph = builder.compile(checkpointer=checkpointer)
config = {"configurable": {"thread_id": "ticket-1001"}}
first = graph.invoke(
{"ticket": "I want a refund for order 55", "category": "", "draft": "", "log": [], "sent": False},
config,
{"recursion_limit": 25},
)
print(first["__interrupt__"][0].value)
# {'draft': 'Thanks for getting in touch about your refund request.', 'category': 'refund'}
final = graph.invoke(Command(resume="approve"), config)
print(final["sent"], final["log"])
# True ['classified as refund', 'drafted a reply', 'human said approve']
Read the file against the sections above and nothing in it should be new. State is a TypedDict with one accumulating key, log, which has a reducer because several nodes append to it. Four nodes, each returning a partial update. One conditional edge with its destinations declared so the diagram is accurate. A SQLite checkpointer, so closing the program and reopening it on thread ticket-1001 picks the ticket up where it stopped. One interrupt with its side effect placed after it rather than before. An explicit recursion_limit.
Two things to try with it that turn it from an example into understanding. First, run only the first half, kill the process, start a fresh one, and resume with Command(resume="approve") — the ticket is still waiting in triage.db. Second, print graph.get_state_history(config) and read the thread backwards: every step, with the state as it was and the writes that produced it.
Then take it further in two directions. Replace classify with a real model call and the hard-coded keyword check disappears. Deploy it: write a langgraph.json listing your graph, and langgraph dev serves it with an API and Studio attached.
{
"dependencies": ["."],
"graphs": { "triage": "./triage.py:graph" },
"env": ".env"
}
- Run
triage.pyend to end, then do the kill-and-resume version. - Send a ticket with no refund in it and confirm it takes the
sendbranch with no pause. - Write
langgraph.json, runlanggraph dev, and step through the graph in Studio.
What you can now do, and what comes next
You can define a state schema and know why a node returns a partial update rather than the whole thing. You can write nodes, wire normal and conditional edges, build a loop and stop it deliberately. You know what a reducer is, why a parallel write without one fails, and why returning [] does not clear an accumulating list. You can attach a checkpointer, give a run a thread ID, inspect any past checkpoint, and resume a graph in a new process. You can stream updates, values and tokens. You can pause for a human and resume safely, and you know the four rules that keep that safe. You can use create_agent when the loop is the standard one, and you can read LangGraph's errors as instructions.
That is genuinely enough to ship something small and useful. The gaps are real, though, and worth naming so you know what you do not yet know.
Persistence and memory in depth. The store for cross-thread memory, semantic search over it, checkpoint TTLs, and what durability mode to choose for a given workload.
Subgraphs. A compiled graph used as a node, which is how a system grows past a dozen nodes without becoming unreadable — along with the three persistence modes a subgraph can have and the MULTIPLE_SUBGRAPHS error.
Map-reduce with Send. Dynamic fan-out from a conditional edge, where you do not know until runtime how many parallel branches there are.
Reliability. RetryPolicy and which exceptions it retries by default, caching with CachePolicy, per-node timeouts, node-level error handlers, and graceful shutdown with RunControl — most of which arrived in 1.2.
Time travel. Replaying from a past checkpoint, and forking a thread with update_state to explore an alternative.
Context and runtime. Passing dependencies with context_schema and reading them through Runtime instead of smuggling them through config.
Deployment. The Agent Server, assistants, crons, double-texting strategies, authentication, and what a real LangSmith Deployment looks like on Postgres and Redis.
Mid-level covers those, plus testing graphs properly and debugging beyond print statements. Senior covers architecture, scaling, multi-tenancy, encryption at rest, migrations against in-flight threads, and when to pick something else entirely.
Two practical next steps while this is fresh. First, instrument what you built: set LANGSMITH_TRACING=true and LANGSMITH_API_KEY, run the triage graph, and look at the trace. Seeing a run as a tree makes the super-step model click in a way that reading about it does not. If you prefer open-source tracing, the Langfuse guide covers the same job with a self-hostable backend, which matters when data residency rules keep traces inside the Gulf or Egypt. Second, when your graph starts retrieving documents, you will want to know whether the answers are any good — that is what the RAGAS guide is for. If you want the higher-level agent framework rather than the runtime, read the LangChain guide next.
One last habit. Pin your versions. langgraph and its checkpointer packages should move together, minor releases land every one to two months and patches often weekly, and a breaking change only arrives in a major version. A lock file is the difference between upgrading when you choose to and upgrading when something breaks.
Sources
- LangGraph overview
- Install LangGraph
- Quickstart
- Graph API
- Use the Graph API
- Choosing between the Graph and Functional APIs
- Pregel runtime
- Persistence
- Checkpointers
- Add memory
- Streaming
- Interrupts and human-in-the-loop
- Durable execution
- Run a local server
- Application structure
- Common errors
- GRAPH_RECURSION_LIMIT
- INVALID_CONCURRENT_GRAPH_UPDATE
- MISSING_CHECKPOINTER
- LangGraph v1 release notes
- Release policy and versioning
- LangSmith CLI reference