This is part one of three. It covers everything you need to build a working LLM application with LangChain, not a teaser. By the end you can talk to a model from Python, give it tools it can call, wrap that into an agent that loops until a task is done, keep a conversation going across turns, force the output into a shape your code can use, and read the errors LangChain throws when you get it wrong. Mid-level and Senior take the same topics further; nothing here is thrown away.
Everything below is written against LangChain 1.4 (version 1.4.3, released 28 September 2026) on Python 3.10 or newer. That matters more for LangChain than for most libraries, because version 1.0 changed the main API, and a large share of the tutorials and blog posts you will find online are still written for 0.x. Each section ends with a Try it task. Do them as you go.
create_react_agent, it is out of date
The modern entry point is from langchain.agents import create_agent. The older from langgraph.prebuilt import create_react_agent is deprecated. Likewise prompt= is now system_prompt=, and LLMChain and friends moved to a separate langchain-classic package. When something you copy does not work, check its date before you doubt yourself.
What LangChain is, and the problem it solves
LangChain is a Python library for building applications on top of large language models. It gives you one consistent way to call any provider's model, one consistent way to describe a function the model is allowed to call, and a ready-made loop that lets the model keep calling those functions until it has finished a task.
To see why that is worth a library, consider what you have without one. A model provider gives you an HTTP API. You send a list of messages and some JSON describing the functions the model may use; you get back a response that either contains text or a request to call one of your functions. If it asks for a function call, you have to parse the arguments, run the function yourself, append the result to the message list in exactly the shape that provider expects, and send the whole thing again. Then you do it again, and again, until the model stops asking. Along the way you handle rate limits, retries, timeouts, and the fact that every provider spells all of this slightly differently.
None of that is hard. All of it is fiddly, and none of it is your product. The first time you write it you learn something. The fourth time, when you are porting it from one provider to another because the pricing changed, you are just paying a tax.
That diagram is the whole idea, and the loop in it is what the documentation calls an agent. The docs frame it as a formula worth memorising: an agent is a model plus a harness. The model decides what to do next. The harness is everything around that decision: the system prompt that sets the model's job, the tools it is allowed to call, and the policies that run before and after each step. LangChain's create_agent is a configurable harness. You supply the model, the tools and the prompt; it supplies the loop.
The other half of LangChain's value is provider independence. A chat model in LangChain is an object with a standard interface, and the provider-specific code lives in a separate small package: langchain-openai, langchain-anthropic, and so on. Swapping "openai:gpt-5.5" for "anthropic:claude-sonnet-4-6" is a one-string change. Because those packages pass model names straight through to the provider's API, a new model usually works the day it is released without upgrading LangChain at all.
- Write down, in plain English, one task you would like a model to do that requires looking something up — a database query, a file read, an HTTP call.
- List the functions the model would need permission to call to do it.
- Keep that list. It is the
tools=argument you will write later in this guide.
The four nouns you need: model, message, tool, agent
Almost everything in beginner LangChain is one of four things. Learn these four and the API stops feeling large.
A chat model is an object that takes a list of messages and returns one message. Its class name starts with Chat — ChatOpenAI, ChatAnthropic — and the base class is BaseChatModel. You rarely need the class itself, because init_chat_model("provider:model") builds the right one for you. There is an older family of "LLM" classes that take a string and return a string; if you ever find yourself holding a plain str where you expected a message object, you picked one of those by mistake.
A message is one turn in the conversation, and there are four types you will meet immediately. A SystemMessage carries standing instructions. A HumanMessage is what the user said. An AIMessage is what the model said, and it may carry tool_calls — requests to run your functions. A ToolMessage carries the result of one of those calls back to the model, tagged with the tool_call_id it answers. You can write messages as plain dictionaries, {"role": "user", "content": "..."}, which is what most of this guide does because it is shorter.
A tool is a schema paired with a function: a name, a description, and a typed argument list, plus the Python callable that actually runs. The @tool decorator builds the schema from your function's type hints and docstring, which is why both are mandatory rather than polite.
An agent is the loop. create_agent(model=..., tools=..., system_prompt=...) returns a compiled graph object. You call .invoke() on it with a list of messages, and it runs the model, executes any tools the model asked for, feeds the results back, and repeats until the model answers without asking for a tool.
One extra fact explains a lot of LangChain's behaviour: create_agent returns a compiled LangGraph graph. LangGraph is the lower-level orchestration library underneath, and LangChain depends on it. You do not need to learn LangGraph to use create_agent, but you inherit its features for free — streaming, persistence across turns, pausing for human approval — and you also inherit its vocabulary in error messages. When an exception mentions a "super-step" or a "recursion limit", that is LangGraph talking. Our LangGraph guide is the natural next stop once you want to control the loop yourself rather than configure it.
- For each of the four nouns, write one sentence from memory: what it holds and what it returns.
- Then answer: which of the four can carry a
tool_call_id, and why does it need one?
ToolMessage, because the model may have asked for several tool calls in one turn and the IDs are how each result is matched to its request. If you can say that sentence, the hardest part of the message model is already behind you.
Installing LangChain and checking that it works
LangChain is pure Python, so installation is identical on Linux, macOS and Windows. You need Python 3.10 or newer — 3.9 was dropped in version 1.0 — and 3.11 or newer if you later want the LangGraph CLI.
Always install into a virtual environment. LangChain pulls in a graph of pinned dependencies (langchain 1.4.3 requires langchain-core>=1.6.3 and langgraph>=1.2.11,<1.3), and letting that mix with your system Python is how you end up with a version conflict that is hard to unpick.
python3 -m venv .venv && source .venv/bin/activate
pip install -U "langchain[openai]"
On Windows PowerShell the first two lines differ and the rest is the same:
py -m venv .venv
.venv\Scripts\Activate.ps1
pip install -U "langchain[openai]"
The square brackets are an extra: an optional group of dependencies. langchain[openai] installs langchain plus langchain-openai. The quotes matter — zsh, which is the default shell on macOS, treats an unquoted [openai] as a filename pattern and fails before pip ever runs. The available extras include openai, anthropic, google-genai, google-vertexai, azure-ai, aws, ollama, mistralai, groq, huggingface, deepseek, together, xai and mcp. Installing the provider package directly (pip install langchain-openai) does exactly the same thing.
Now give it a key. Models are remote services and every provider wants credentials, read from an environment variable:
export OPENAI_API_KEY="sk-..." # or ANTHROPIC_API_KEY, GOOGLE_API_KEY
For Azure OpenAI you need three: AZURE_OPENAI_API_KEY, AZURE_OPENAI_ENDPOINT and OPENAI_API_VERSION. For Bedrock, your normal AWS credentials. Set these in your shell profile or a .env file you never commit, not in your source.
Check the installation in two steps. First confirm the versions, then confirm the imports:
pip show langchain langchain-core langgraph
python -c "import langchain; print(langchain.__version__)"
python -c "from langchain.agents import create_agent; from langchain.chat_models import init_chat_model; print('ok')"
You should see 1.4.3 or later from the second command and ok from the third. If the import fails, you almost certainly have an old major version in the environment; pip install -U langchain and check again.
There is one more check worth running, and it is the most useful of the four because it needs no API key and costs nothing. LangChain ships a fake chat model for tests that returns canned replies:
from langchain_core.language_models.fake_chat_models import GenericFakeChatModel
from langchain.agents import create_agent
agent = create_agent(GenericFakeChatModel(messages=iter(["hello"])), tools=[])
result = agent.invoke({"messages": [{"role": "user", "content": "hi"}]})
print(result["messages"][-1].content)
Run it and you get hello. That one script proves your Python version, your install, and the whole agent loop are working, with no network call and no bill. Keep it: the same fake model is how you will write unit tests later.
- Create a virtual environment and install
"langchain[openai]"or"langchain[anthropic]". - Run all three verification commands and the
verify_setup.pyscript. - Deliberately break it: run
python -c "from langchain.chains import LLMChain".
ModuleNotFoundError: No module named 'langchain.chains'. That is the single most common error people hit after following an old tutorial, and now you have seen it on purpose. The fix, if you ever genuinely need those legacy chains, is pip install langchain-classic and importing from langchain_classic instead.
Your first conversation with a model
Start with the model alone, before any agent. Everything else in LangChain is built on this one object, and understanding what it returns saves hours later.
from langchain.chat_models import init_chat_model
model = init_chat_model("openai:gpt-5.5")
response = model.invoke("Why is the sky blue?")
print(type(response))
print(response.text)
Three things to notice. init_chat_model takes a "provider:model" string; it can often infer the provider from the model name alone, and model_provider= sets it explicitly when you want no ambiguity. invoke accepts a bare string as a shorthand for a single human message. And the return value is not a string — it is an AIMessage. The text lives in response.text.
response.text is a property, not a method. In 0.x it was message.text(). The call form still works and emits a warning, and it is scheduled for removal in version 2, so write it without the parentheses. This is the second-most-common symptom of copying an old example.
An AIMessage carries more than text, and the extras are what make it worth having:
from langchain.chat_models import init_chat_model
model = init_chat_model("openai:gpt-5.5")
response = model.invoke("Name three cities in Egypt.")
print(response.text) # the answer as a string
print(response.usage_metadata) # input/output token counts
print(response.response_metadata) # provider's raw extras: model name, finish reason
print(response.content_blocks) # typed, provider-agnostic view of the content
print(response.tool_calls) # [] here: the model asked for nothing
usage_metadata is how you find out what a call cost. content_blocks is newer and solves a real problem: modern models return mixed content — text, reasoning traces, images, citations, tool calls — and every provider shapes that differently. content_blocks presents it as a list of typed blocks (text, reasoning, image, audio, file, tool_call and others) with the same shape whichever provider produced it. The raw content attribute is still there, unchanged, when you want exactly what the provider sent.
You can configure the model when you build it. The parameters are standard across providers:
from langchain.chat_models import init_chat_model
model = init_chat_model(
"anthropic:claude-sonnet-4-6",
temperature=0.7, # higher = more varied output
max_tokens=1000, # cap on the response length
timeout=30, # seconds before giving up on one request
max_retries=6, # the default
)
max_retries deserves attention because it already solves a problem you might otherwise write code for. The default is 6, with exponential backoff and jitter, and it retries network errors, HTTP 429 rate limits and 5xx server errors. It does not retry a 401 or a 404, which is correct: a wrong API key or a misspelled model name will not fix itself no matter how many times you ask.
Two other ways of calling the model are worth knowing on day one. Streaming prints tokens as they arrive instead of waiting for the whole answer, which is the difference between an interface that feels broken and one that feels fast:
from langchain.chat_models import init_chat_model
model = init_chat_model("openai:gpt-5.5")
for chunk in model.stream("Write two sentences about the Nile."):
print(chunk.text, end="", flush=True)
print()
And batching sends several independent prompts with controlled concurrency:
from langchain.chat_models import init_chat_model
model = init_chat_model("openai:gpt-5.5")
answers = model.batch(
["Capital of Jordan?", "Capital of Oman?", "Capital of Tunisia?"],
config={"max_concurrency": 5},
)
for a in answers:
print(a.text)
batch preserves input order. batch_as_completed yields results as they finish, out of order, which is better when you want to start processing early. Note that this is client-side concurrency — LangChain firing several HTTP requests — and not the same thing as a provider's own cheaper batch API.
- Run
first_call.py, then printresponse.usage_metadataand note the token counts. - Change the model string to a different provider you have a key for. Change nothing else.
- Run the streaming example, then run it again with
temperature=0and compare the two answers.
temperature=0 repeated runs are near-identical, which is what you want when you are testing.
Messages: how a conversation is actually represented
A single prompt string is a convenience. Real applications pass a list of messages, because that list is the model's memory. A chat model is stateless: it knows nothing except what is in the list you send.
from langchain.chat_models import init_chat_model
from langchain.messages import SystemMessage, HumanMessage
model = init_chat_model("openai:gpt-5.5")
conversation = [
SystemMessage("You are a terse assistant. Answer in one sentence."),
HumanMessage("What is a vector database?"),
]
first = model.invoke(conversation)
print(first.text)
conversation.append(first)
conversation.append(HumanMessage("Name one."))
second = model.invoke(conversation)
print(second.text)
Read the second call carefully. The model answers "Name one" correctly only because the previous question and answer are in the list. Drop the append lines and it has no idea what "one" refers to. There is no hidden session. When you later add a checkpointer, all it does is keep this list for you between calls.
The same conversation in dictionary form, which is shorter and equally valid:
conversation = [
{"role": "system", "content": "You are a terse assistant. Answer in one sentence."},
{"role": "user", "content": "What is a vector database?"},
]
Accepted roles are system, user, assistant and tool. LangChain also accepts (role, content) tuples and a bare string. What it does not accept is a dictionary with the wrong keys, and the error it raises is unusually clear:
ValueError: Message dict must contain 'role' and 'content' keys,
got {'role': 'HumanMessage', 'random_field': 'random value'}
That is the MESSAGE_COERCION_FAILURE error. The usual cause is putting a class name where a role belongs — {"role": "HumanMessage", ...} instead of {"role": "user", ...}.
The SystemMessage is the one with outsized influence. It sets the model's job, its tone, its constraints and anything it should refuse. It is not magic and a determined user can argue with it, but it is the cheapest quality lever you have: most "the model keeps doing the wrong thing" problems at this level are solved by writing three more specific sentences in the system prompt.
.txt or .md file loaded with Path("prompts/agent.md").read_text() is enough at this stage.
- Run
messages.pyas written, then delete the twoappendlines and run it again. - Reproduce the coercion error on purpose by sending
{"role": "Human", "content": "hi"}. - Rewrite the system message to demand answers only in bullet points, and confirm the change.
Writing your first tool
A tool is how the model reaches the world outside its own weights. The model never runs your code — it emits a request to run it, and the harness does the running. That separation is why tools are also the main place security matters, a theme Senior level develops.
from langchain.tools import tool
@tool
def get_shipping_cost(country: str, weight_kg: float) -> str:
"""Calculate the shipping cost for a parcel to a given country.
Args:
country: Destination country name, e.g. "Egypt"
weight_kg: Parcel weight in kilograms
"""
rates = {"egypt": 4.0, "jordan": 5.5, "oman": 6.0}
rate = rates.get(country.lower())
if rate is None:
return f"No shipping rate configured for {country}."
return f"{rate * weight_kg:.2f} USD"
Every line of that is load-bearing. The function name becomes the tool name the model sees, so it should describe the action in snake_case. The type hints become the argument schema; without them LangChain cannot tell the model what to send. The docstring becomes the tool description, and it is the only thing the model reads when deciding whether this tool is relevant — a vague docstring produces a tool the model ignores or misuses. The return value goes back to the model as a ToolMessage, so return something a model can read, not an internal object.
Inspect what the model will actually be shown:
from tools import get_shipping_cost
print(get_shipping_cost.name)
print(get_shipping_cost.description)
print(get_shipping_cost.args_schema.model_json_schema())
print(get_shipping_cost.invoke({"country": "Egypt", "weight_kg": 2.5}))
The last line runs the tool directly, bypassing any model. Do this for every tool you write before you give it to an agent. A tool that raises on its own is a tool that will make your agent look broken for reasons that have nothing to do with the model.
The decorator takes options when you need them: @tool("search") renames it, @tool(description="...") overrides the docstring, @tool(args_schema=MyPydanticModel) replaces the inferred schema with an explicit one, and @tool(return_direct=True) ends the agent loop immediately with the tool's output instead of sending it back to the model.
A tool a model uses well
get_order_status(order_id: str)- Docstring says what it returns and when to use it
- Every argument typed
- Returns a short readable string
- Handles its own bad input and says so
A tool a model misuses
do_thing(data)- Docstring: "helper"
- Untyped
**kwargs - Returns a nested dict of database rows
- Raises
KeyErroron anything unexpected
Two naming rules will save you a debugging session. Use snake_case, because some providers reject tool names containing spaces or special characters. And never name a tool parameter config or runtime: both are reserved by LangChain for injecting framework objects, and using them as ordinary arguments causes a runtime error. If you need access to per-run information inside a tool, the supported way is a parameter annotated runtime: ToolRuntime, which Mid-level covers.
- Write a tool from the list you made in the first section. Give it full type hints and a real docstring.
- Print its
.descriptionand its JSON schema. Read them as if you were the model. - Call
.invoke({...})with valid input, then with input you know is wrong.
Your first agent
Now put the pieces together. This is the shortest useful LangChain program there is.
from langchain.agents import create_agent
from tools import get_shipping_cost
agent = create_agent(
model="openai:gpt-5.5",
tools=[get_shipping_cost],
system_prompt=(
"You are a shipping assistant. Use the shipping tool to quote prices. "
"Never guess a price. If a country has no rate, say so plainly."
),
)
result = agent.invoke({
"messages": [{"role": "user", "content": "What does a 3 kg parcel to Oman cost?"}]
})
print(result["messages"][-1].content)
model= accepts a model string or a model object you built with init_chat_model. tools= is a plain list. system_prompt= is a string, or a SystemMessage from version 1.1 onward.
The return value is a dictionary, and result["messages"] is the full transcript — not just the final answer. Printing all of it is the single best way to understand what an agent did:
from agent import agent
result = agent.invoke({
"messages": [{"role": "user", "content": "What does a 3 kg parcel to Oman cost?"}]
})
for message in result["messages"]:
print(f"--- {type(message).__name__}")
print(message.text)
if getattr(message, "tool_calls", None):
print("asked for:", message.tool_calls)
You will see four messages: your HumanMessage; an AIMessage with empty text and a tool_calls entry naming get_shipping_cost with arguments {"country": "Oman", "weight_kg": 3}; a ToolMessage containing 18.00 USD; and a final AIMessage with the sentence for the user. That is the loop, fully visible. Read this output every time an agent surprises you — in most cases the model either called the wrong tool or was handed a tool result it could not interpret, and both are obvious here.
Three constraints on create_agent exist because of the version 1 redesign, and each has caught experienced users:
You cannot pass a pre-bound model. create_agent(ChatOpenAI().bind_tools([...]), tools=[]) is rejected. The agent binds the tools itself from tools=, and doing it twice would give the model two conflicting tool lists.
You cannot pass a ToolNode in tools=. That was the 0.x way of customising tool error handling; it is now done with middleware.
If you add custom state, it must be a TypedDict subclassing langchain.agents.AgentState. Pydantic models and dataclasses are no longer accepted as agent state. (Dataclasses are still the normal choice for the separate context_schema, which is immutable per-run data rather than state.)
One behavioural detail worth knowing before you need it: the loop is bounded. LangGraph counts "super-steps" and stops at recursion_limit, which defaults to 1000 in current versions. Tutorials written against 0.x say 25. If an agent ever runs away, this is the backstop that eventually raises GraphRecursionError, and the fix is almost never to raise the limit.
- Build the agent with your own tool and ask it a question that requires the tool.
- Print the whole transcript with
trace_agent.pyand identify all four message types. - Now ask it something unrelated, like "what is the capital of Morocco?", and print the transcript again.
Remembering the conversation: threads and checkpointers
Run the agent twice and the second call knows nothing about the first. Each invoke starts from the messages you passed. To hold a conversation you need two things: a checkpointer that saves state, and a thread ID that says which conversation you mean.
from langchain.agents import create_agent
from langgraph.checkpoint.memory import InMemorySaver
agent = create_agent(
model="openai:gpt-5.5",
tools=[],
system_prompt="You are a helpful assistant.",
checkpointer=InMemorySaver(),
)
config = {"configurable": {"thread_id": "customer-42"}}
agent.invoke({"messages": [{"role": "user", "content": "My name is Layla."}]}, config=config)
second = agent.invoke({"messages": [{"role": "user", "content": "What is my name?"}]}, config=config)
print(second["messages"][-1].content)
It answers "Layla". Change thread_id to something else and it does not, because that is a different conversation. Notice that the second call passes only the new message — the checkpointer supplies the history, so you do not rebuild the list yourself.
InMemorySaver keeps everything in the Python process, which means it vanishes when the process exits and is invisible to any other process. It is right for development and tests and wrong for anything else. The durable options are SqliteSaver for a local file (pip install langgraph-checkpoint-sqlite) and PostgresSaver for production (pip install -U langgraph-checkpoint-postgres "psycopg[binary]"). The code change is one line, which is the point of learning the pattern now.
A checkpointer is doing more than remembering chat history. It persists the whole graph state at every super-step, and that is what makes several later features possible: resuming a run that was paused for human approval, inspecting or rewinding to an earlier state, and surviving a crash mid-run.
Which leads to an error you will certainly meet:
MISSING_CHECKPOINTER
It is raised when you use a feature that requires persistence without providing one — a thread_id, an interrupt for human approval, agent.get_state(), or thread-scoped call limits. The fix is to pass checkpointer=. There is one important exception: if you later deploy to LangSmith's managed Agent Server, a checkpointer is provisioned for you and you should not pass your own.
Short-term memory inside a thread is distinct from long-term memory across threads. The thread holds this conversation. For facts that should outlive it — a user's preferences, say — LangGraph offers a store, a JSON document store keyed by a namespace tuple and a key, with put, get and search. Mid-level covers it; for now, know that the distinction exists so you do not try to make a thread do that job.
ContextOverflowError. The real fixes — summarising old turns, trimming history, clearing stale tool output — are middleware features covered at Mid-level. For now just know the cost is linear in conversation length, which is why a long-running chat gets slower and more expensive as it goes.
- Run
memory.py, then change thethread_idand confirm the agent forgets. - Remove
checkpointer=but keep theconfig, and read the error you get. - Swap
InMemorySaverforSqliteSaver, run the script, exit Python, and run only the second question.
Getting structured output your code can use
Prose is fine for a human and awkward for a program. When the next step is a database insert or an if statement, you want a typed object. LangChain gives you that with response_format=.
from pydantic import BaseModel, Field
from langchain.agents import create_agent
class Ticket(BaseModel):
category: str = Field(description="One of: billing, technical, account")
urgency: int = Field(description="1 (low) to 5 (critical)")
summary: str = Field(description="One sentence summarising the issue")
agent = create_agent(
model="openai:gpt-5.5",
tools=[],
system_prompt="Classify incoming support tickets.",
response_format=Ticket,
)
result = agent.invoke({"messages": [
{"role": "user", "content": "I was charged twice this month and support never replied."}
]})
ticket = result["structured_response"]
print(ticket.category, ticket.urgency)
print(type(ticket))
The parsed object lands under result["structured_response"], not in the last message, and it is a real Ticket instance with validated fields. The Field(description=...) text is sent to the model as part of the schema, so it is prompt engineering with a different syntax — "1 (low) to 5 (critical)" does real work there.
Outside an agent, use the model directly:
from langchain.chat_models import init_chat_model
from structured import Ticket
model = init_chat_model("openai:gpt-5.5").with_structured_output(Ticket)
print(model.invoke("The app crashes when I open settings.").category)
Under the hood there are two strategies, and response_format=Ticket picks one for you. ToolStrategy asks the model to "call" an artificial tool whose arguments are your schema, which works on any model that supports tool calling. ProviderStrategy uses the provider's own native structured-output feature, which is more reliable where it exists. You can name the strategy explicitly when you need control:
from langchain.agents.structured_output import ToolStrategy, ProviderStrategy
from structured import Ticket
explicit_tool = ToolStrategy(Ticket)
explicit_native = ProviderStrategy(Ticket)
Pydantic is not the only accepted schema. A dataclass or TypedDict works and returns a dict rather than an instance, and a raw JSON Schema dictionary works but must be wrapped in an explicit strategy and must carry a top-level title and description.
One thing that was removed in version 1 and still appears in old examples: prompted output. You could once write response_format=("please generate JSON like...", Schema). That no longer works. If you find it in a tutorial, the tutorial predates version 1.
- Define a Pydantic schema for something in your own domain and classify three inputs with it.
- Delete all the
Field(description=...)text and run the same three inputs again. - Add a field the input text cannot possibly support, such as
customer_tier: str, and see what the model invents.
Watching what happens: streaming and tracing
Two kinds of visibility matter from your first week. Streaming shows your user that something is happening. Tracing shows you what happened.
Streaming an agent is not quite the same as streaming a model, because an agent produces several kinds of event: model tokens, tool calls, state updates. stream_mode selects which you receive.
from agent import agent
for token, metadata in agent.stream(
{"messages": [{"role": "user", "content": "What does a 2 kg parcel to Jordan cost?"}]},
stream_mode="messages",
):
print(token.text, end="", flush=True)
print()
stream_mode="messages" yields (token, metadata) tuples, where the metadata tells you which node produced the token. stream_mode="updates" yields the state change after each step, which is better for debugging than for display. You can pass a list of modes, and with the default stream format each chunk then arrives as a (mode, chunk) tuple.
from agent import agent
for update in agent.stream(
{"messages": [{"role": "user", "content": "Quote 1 kg to Egypt."}]},
stream_mode="updates",
):
print(update)
You will see a dictionary keyed by node name — model, then tools, then model again. That node is called "model". In 0.x it was called "agent", and code that filters on "agent" silently matches nothing.
Tracing is the bigger win, and it needs no code at all. LangSmith is LangChain's hosted observability platform, and two environment variables turn it on:
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY="lsv2_..."
export LANGSMITH_PROJECT="shipping-assistant" # optional; defaults to "default"
Run your agent again and every call appears in the LangSmith UI as a tree: the agent run, each model call with its exact prompt and response, each tool call with its arguments and result, token counts and latency at every level. The first time you look at a trace of an agent that misbehaved, the cause is usually obvious within seconds — a tool description the model misread, or a tool returning an error string the model treated as data.
LANGCHAIN_TRACING_V2 and LANGCHAIN_API_KEY in older posts and even a few doc pages. The documented current names are the LANGSMITH_* ones. Use those.
If a hosted service is not an option — and for some employers in the region it is not — you can get useful visibility locally by printing the full transcript as in trace_agent.py, and by passing config={"run_name": ..., "tags": [...], "metadata": {...}} to invoke, which labels runs for whatever tracer you do use. Our Langfuse guide and LangSmith guide go further into the options.
- Stream your agent with
stream_mode="messages"and watch the answer appear token by token. - Switch to
stream_mode="updates"and note the node names in the output. - If you can, create a free LangSmith key, set the two variables, and open the trace of one agent run.
model and tools. In a LangSmith trace, find the exact JSON arguments the model sent to your tool — that single view answers most "why did it do that?" questions.
Configuration and the errors you will actually hit
LangChain's configuration surface at this level is small: environment variables for credentials and tracing, constructor arguments on the model, and a config dictionary on each call. The config dictionary carries configurable values such as thread_id, plus run_name, tags, metadata, callbacks, max_concurrency and recursion_limit. Tags and metadata are inherited by child runs; run_name is not.
Errors are where a beginner spends real time, so learn to read them. LangChain exceptions carry an lc_error_code attribute matching a documented error page, which means the error itself tells you what to search for.
| Error code or exception | What actually happened | What to do |
|---|---|---|
ModuleNotFoundError: langchain.chains |
You followed a pre-1.0 tutorial | pip install langchain-classic, import from langchain_classic |
MESSAGE_COERCION_FAILURE |
A message dict is missing role or content |
Use {"role": "user", "content": "..."} or a message class |
MISSING_CHECKPOINTER |
Used thread_id, an interrupt or get_state with no checkpointer |
Pass checkpointer=InMemorySaver() |
ModelAuthenticationError |
Key missing, wrong, or never loaded | Check the variable name and that .env was loaded; try api_key= explicitly |
ModelNotFoundError |
Model ID typo, or not enabled on your account | Fix the string; check your provider console |
ModelRateLimitError |
Provider throttled you | It is retryable and max_retries=6 already backs off; add a rate_limiter if persistent |
ContextOverflowError |
Input exceeds the model's context window | Shorten the input or trim the conversation |
GraphRecursionError |
More super-steps than recursion_limit (default 1000) |
Find the loop; do not just raise the limit |
INVALID_CHAT_HISTORY |
An AIMessage with tool_calls has no matching ToolMessage |
Supply one ToolMessage per call, with matching tool_call_id |
The response is a str, not an AIMessage |
You used a legacy completion class | Use a Chat* class via init_chat_model |
LangChainBetaWarning on import langchain.mcp |
Expected: that namespace is beta in 1.4 | Nothing; filter it with warnings if it bothers you |
Three of those deserve a sentence more. The model exceptions in langchain_core.exceptions each carry an is_retryable flag, and the division is sensible: authentication, permission, invalid request, not-found and context-overflow are not retryable; rate limit, API error, connection error and timeout are. If you write your own retry logic, check that flag rather than guessing from the message.
INVALID_CHAT_HISTORY is the one that looks mysterious. It happens when you send a conversation where the model asked for a tool and never got an answer — typically because you caught an exception from a tool and carried on, or because you resumed a paused run incorrectly. Providers enforce this rule strictly: every tool call needs exactly one matching result, no duplicates and no orphans.
ModelAuthenticationError has an unglamorous most-common cause: the .env file was never loaded. LangChain reads os.environ; it does not read .env for you. Either export the variables in your shell or call load_dotenv() yourself.
- Cause three errors on purpose: a misspelled model name, a
thread_idwith no checkpointer, and an unset API key. - For each, print
type(exc).__name__andgetattr(exc, "lc_error_code", None). - Write the three codes into your notes next to what you did to trigger them.
Choosing a provider, and the packages around LangChain
Beginners are often confused by how many packages have "langchain" in the name, and by which one a given import comes from. The structure is simpler than it looks, and knowing it turns a confusing ImportError into a one-line fix.
langchain-core holds the abstractions and nothing else: the message classes, the tool interface, the chat model base class, the document and vector store interfaces. It has almost no dependencies, which is deliberate — provider packages depend on it without dragging anything else in. Occasionally you import from it directly, as this guide did for GenericFakeChatModel and as you would for Document or UsageMetadataCallbackHandler.
langchain is the package you install. In version 1 its namespace was deliberately shrunk to the things a modern application needs: agents, messages, tools, chat_models, embeddings, plus smaller modules such as rate_limiters and the new mcp. If an import from langchain.something fails, the first question is whether that module ever existed in version 1.
langgraph is the runtime underneath. create_agent builds a LangGraph graph, so checkpointers, interrupts and streaming all come from here — which is why from langgraph.checkpoint.memory import InMemorySaver is a langgraph import in otherwise pure LangChain code.
Provider packages are one per provider: langchain-openai, langchain-anthropic, langchain-google-genai, langchain-ollama, and so on. They are versioned independently and released frequently, which is how a model that launched this morning works this afternoon.
langchain-classic holds what version 1 moved out: LLMChain, ConversationChain, the legacy retriever classes, the indexing API, and the prompt hub. It is a compatibility package, not a deprecated one in the sense of broken — but if you are writing new code and find yourself reaching for it, there is usually a version 1 way to do the same thing. langchain-community is a large grab-bag of third-party integrations that is being sunset; prefer a dedicated provider package where one exists. langchain-experimental is archived and should not be used.
Which provider should you actually pick? For learning, any one you can get a key for. For work, the question is rarely about raw model quality and usually about three things: where the data may go, what the per-token cost is at your volume, and what your organisation already has a contract for. In practice that means teams in Egypt and the Gulf often land on Azure OpenAI (because the data residency story and the procurement path are already solved), on Bedrock for AWS-centric shops, or on a self-hosted open-weights model served through Ollama or vLLM for the strictest cases. All of those are the same init_chat_model call with different arguments.
Two configuration details let LangChain reach almost anything. base_url= points an OpenAI-compatible client at a different endpoint, which is how self-hosted servers such as vLLM are used; pass model_provider="openai" alongside it so the right client is chosen. And a rate_limiter built with InMemoryRateLimiter from langchain.rate_limiters throttles your own request rate, which is worth knowing exists before a shared development key gets you throttled at the worst moment.
- Run
pip list | grep langchainand name what each installed package is for. - For each import used in this guide so far, say which package it came from without looking.
- Check whether your organisation or university already has access to a model endpoint, and what region it runs in.
langchain.* for agents, tools, messages and chat models; langgraph.* for the checkpointer; langchain_core.* for the test fake. If you can place each one, package errors stop being mysterious.
What one tool call looks like from the model's side
It is worth slowing down on the single most important mechanism in the whole guide, because almost every confusing agent behaviour traces back to a misunderstanding of it.
When you pass tools=[get_shipping_cost] to create_agent, LangChain converts that function into a JSON schema and sends it to the provider alongside your messages. The model never sees your Python. It sees a name, a description, and an argument schema — exactly the three things you printed with inspect_tool.py. That is the entire basis on which it decides whether and how to use your tool, which is why the docstring is not documentation but input.
The model then does one of two things. It either answers in text, or it returns an AIMessage whose content is often empty and whose tool_calls list contains one or more requests, each with a tool name, an argument dictionary and a generated id. Crucially, the model has not run anything and cannot: it has produced structured text asking for a function to be called.
The harness takes over. For each requested call it finds the matching tool, validates the arguments against the schema, runs the function, and wraps the return value in a ToolMessage carrying the same id in its tool_call_id. Those messages are appended to the conversation and the whole thing goes back to the model, which now sees its own request and the result next to each other.
Four consequences fall out of that mechanism, and each explains a class of real problem.
The model can ask for several tools at once. A question like "quote 2 kg to Egypt and 3 kg to Oman" may produce two tool_calls in one AIMessage. Every one of them needs exactly one ToolMessage with a matching ID, which is precisely what INVALID_CHAT_HISTORY and INVALID_TOOL_RESULTS are complaining about when they fire.
Tool arguments are model output, so they can be wrong. A model may send weight_kg="three" or a country you have never heard of. Schema validation catches type mismatches; it cannot catch a plausible-looking wrong value. Validate inside the tool and return an explanation the model can act on.
Whatever your tool returns is what the model believes. If your tool returns "error: connection refused", the model treats that as the state of the world and will tell the user about a connection problem. If it returns an empty string, the model will often guess. Return text that makes the situation unambiguous.
A tool that can do damage will be asked to. The security guidance in the official docs is blunt about this: assume the model will eventually use every permission you give it. At beginner level the right instinct is to give tools the narrowest access that works — read-only database credentials, one directory rather than a filesystem, no delete where an update will do. The machinery for requiring human approval before a sensitive tool runs exists (HumanInTheLoopMiddleware, which needs a checkpointer) and is covered at Mid-level, but the habit starts now.
get_user's docstring says "update a user's record", the model will call it to update records and then be confused by the result. When an agent repeatedly does something inexplicable, read your tool descriptions aloud before you touch the system prompt.
- Ask your shipping agent one question that needs two separate quotes, and count the
tool_callsin the transcript. - Make your tool return the single word
"error"and ask a normal question. Read what the agent tells the user. - Now return
"No rate configured for that country. Available: Egypt, Jordan, Oman."and ask again.
"error" produces a vague, unhelpful reply; the explanatory string produces a useful one, often with the valid options offered to the user. Your tool's return text is prompt engineering.
Putting it all together
One project that uses every idea in this guide: a small research assistant with two tools, conversation memory, structured output, and a streamed reply. Create a file and run it.
from dataclasses import dataclass
from pydantic import BaseModel, Field
from langchain.agents import create_agent
from langchain.tools import tool
from langgraph.checkpoint.memory import InMemorySaver
NOTES = {
"cairo-office": "Cairo office opened 2019, 42 staff, lead: Mona Fahmy.",
"riyadh-office": "Riyadh office opened 2023, 11 staff, lead: Khalid Otaibi.",
}
@tool
def lookup_note(key: str) -> str:
"""Look up an internal note by its key.
Args:
key: Note identifier, e.g. "cairo-office"
"""
return NOTES.get(key, f"No note found for key '{key}'. Known keys: {', '.join(NOTES)}")
@tool
def headcount_total() -> str:
"""Return the total headcount across all offices."""
return "53 staff across 2 offices."
class Briefing(BaseModel):
answer: str = Field(description="The answer, in at most two sentences")
sources: list[str] = Field(description="Note keys consulted, empty if none")
agent = create_agent(
model="openai:gpt-5.5",
tools=[lookup_note, headcount_total],
system_prompt=(
"You are an internal briefing assistant. Use lookup_note for office facts "
"and headcount_total for staff totals. Never invent a fact that is not in a "
"note. List every note key you consulted in sources."
),
response_format=Briefing,
checkpointer=InMemorySaver(),
)
config = {"configurable": {"thread_id": "demo-1"}}
def ask(question: str) -> None:
print(f"\n> {question}")
result = agent.invoke({"messages": [{"role": "user", "content": question}]}, config=config)
briefing = result["structured_response"]
print(briefing.answer)
print("sources:", briefing.sources or "none")
if __name__ == "__main__":
ask("Who leads the Cairo office?")
ask("And how many people work there?")
ask("What is our total headcount?")
ask("Who leads the Beirut office?")
Run it and read all four answers, because each one demonstrates something different.
The first question makes the model call lookup_note("cairo-office") and report Mona Fahmy, with that key in sources. The second says only "there" — it works because the checkpointer kept the thread, and the model resolved the pronoun from history. The third needs a different tool, and the model picks headcount_total because the docstrings distinguish them clearly. The fourth is the important one: there is no Beirut note, the tool returns a message saying so along with the known keys, and the model should tell you plainly instead of inventing an answer. If it does invent one, the fix is in the system prompt, not the code.
Then make three changes and watch the behaviour move:
# 1. Delete `response_format=Briefing`. The answer moves back into
# result["messages"][-1].content and there is no sources list to check.
# 2. Delete `checkpointer=InMemorySaver()` and the config. Question two breaks.
# 3. Weaken the system prompt to "You are a helpful assistant."
# Watch the fourth question start producing a plausible fake name.
The third experiment is the one to sit with. Nothing in your Python changed, no tool behaved differently, and the application got worse. The prompt is part of the system, and at this level it is the part you will tune most.
- Run
assistant.pyand confirm all four behaviours. - Make each of the three changes in turn, in isolation, and note what breaks.
- Replace the two toy tools with tools that hit something real — a local JSON file, a SQLite table, a public HTTP API — and keep everything else.
What you can now do, and what comes next
You can install LangChain and verify it without spending money, call any provider's chat model through one interface, read an AIMessage including its token usage and content blocks, build a conversation out of messages and explain why the model has no memory of its own, write tools the model uses correctly, assemble an agent and read its full transcript, persist a conversation with a checkpointer and a thread ID, force output into a Pydantic schema, stream tokens, switch on tracing, and recognise the dozen errors that account for most beginner time lost.
You can also now tell current LangChain from obsolete LangChain, which is a more valuable skill here than in most ecosystems. create_agent not create_react_agent, system_prompt not prompt, .text not .text(), langchain_classic for the old chains, node name "model" not "agent".
What you have not met yet is everything that makes an agent behave well under real conditions. Middleware is the big one: a composable layer that hooks into the loop to summarise long conversations, redact personal data, require human approval before a destructive tool runs, cap the number of model or tool calls, retry intelligently, and pick a cheaper model when the task is simple. ToolRuntime lets a tool read per-run context, write to long-term memory, and emit progress. Context and custom state let you pass a user ID through the whole run without threading it through every function. Retrieval turns a document collection into a tool, which is how RAG is built in version 1.
Mid-level takes all of that, plus event streaming, rate limiting and token accounting, testing agents with fake models, and running an agent as a service with the LangGraph CLI. Senior covers custom middleware, structured-output error handling, multi-agent composition, durability and deployment, the security model, cost control and upgrades.
Three good next guides: LangGraph if you want to control the loop explicitly rather than configure it, LangSmith or Langfuse for observability you can run in your own environment, and RAGAS once you need to measure whether your answers are actually any good. If retrieval is your next step, pick a vector store from Qdrant, pgvector or Chroma and expose it to the agent as a tool.
Before moving on, do one thing: rebuild assistant.py from memory, without looking. The gaps you find are the parts you have read but not yet learned.
Sources
- LangChain Python overview
- Install
- Quickstart
- Agents
- Models
- Messages
- Tools
- Structured output
- Short-term memory
- Streaming
- Observability
- Unit testing
- Common errors
- MISSING_CHECKPOINTER
- MESSAGE_COERCION_FAILURE
- GRAPH_RECURSION_LIMIT
- INVALID_CHAT_HISTORY
- Migrating to LangChain v1
- LangChain v1 release notes
- Changelog
- Release policy
- Checkpointers
- LangSmith environment variables
- langchain on PyPI