This is part one of three. It covers everything you need to do real work with LangSmith, not a teaser. By the end you can send traces from a Python or TypeScript program to LangSmith, open one and read what your application actually did, group a conversation into a thread, attach scores to runs, build a small dataset, run an evaluation over it, and version a prompt. Mid-level and Senior take the same topics further; nothing here is thrown away.
LangSmith is a hosted product from LangChain, so this guide has two jobs. The first is to teach the concepts, which are few and which transfer to every other observability tool you will meet. The second is to be precise about a product that moves quickly. LangSmith has no single version number. The Python SDK, the JavaScript SDK, the command-line tool and the self-hosted Helm chart each have their own, and this guide was checked against Python langsmith 0.14.2, JavaScript langsmith 0.10.5 and langsmith-cli v0.2.60. The SDK ships several releases a week, so run pip install -U langsmith instead of memorising a number.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and tracing only makes sense once you have watched your own program appear in the interface.
What LangSmith is, and the problem it solves
A program that calls a large language model is hard to debug for a reason that has nothing to do with the model being clever. An ordinary function gives the same output for the same input, so a failing test points at a line. An LLM application is a chain of steps: a prompt is assembled, a model is called, the answer is parsed, a tool may be invoked, a retrieval may fetch documents, and the model is called again. Any step can be the one that went wrong, the model is not deterministic, and the user only ever sees the final sentence. When the answer is bad you are left guessing which step produced it.
LangSmith is a platform that records what happened at every step and lets you look at it afterwards. It calls this observability, and it builds two more things on top of the recording: evaluation, which is running your application over a fixed set of examples and scoring the results, and prompt management, which is storing and versioning the prompt templates your application uses.
The diagram is the whole product in one line. Your code is instrumented, a small client library sends what happened to LangSmith without slowing the program down, and you read it there.
Before tools like this existed, people debugged LLM programs with print statements and log files. That works for one request on your laptop. It stops working when a colleague says "the bot gave a strange answer yesterday afternoon", because the prints are gone, the log lines from one request are interleaved with a hundred others, and nobody kept the exact prompt that was sent. A trace is a structured record that keeps the inputs, the outputs, the timings, the token counts and the errors of every step together, tied to one request, and searchable.
You do not need to use LangChain to use LangSmith. The name is confusing because the two come from the same company, and LangChain and LangGraph do trace into LangSmith with no code changes, but LangSmith works with plain Python, plain TypeScript, the OpenAI, Anthropic and Gemini client libraries, and many other frameworks. If you want the framework side, the LangGraph guide is the natural companion to this one, and LangSmith has open-source alternatives such as Langfuse, covered in the Langfuse guide. Comparing the two is a good way to learn what is essential and what is product-specific.
Debug a bad answer
Open the one request that went wrong and see every prompt, tool call and retrieval inside it.
Watch cost and speed
Token counts, latency and errors are recorded per step, so you can see which step is slow or expensive.
Test before you ship
Run a new version of your app over saved examples and compare the scores to the old version.
Version your prompts
Keep prompt templates outside the code, with history, and pull a specific version at runtime.
A note on names, because older tutorials will confuse you. The interface and the docs have been reorganised. The "tenant" of old material is now called a workspace, although the API still uses the words tenant and X-Tenant-Id. Prompt "repos" are now just prompts. The documentation moved from docs.smith.langchain.com, which now redirects, to docs.langchain.com/langsmith. When a tutorial and this guide disagree, trust the current docs.
- Think of an LLM feature you have used that once gave a wrong answer.
- List the steps that probably ran behind it: prompt, model, tools, retrieval, parsing.
- For each step, write down what you would need to see to decide whether that step was at fault.
The vocabulary: runs, traces, projects and threads
LangSmith has a handful of nouns, and every screen and every API call is built from them. Learn them now and the rest of the guide is easy to follow.
A run is one unit of work: a single model call, a single tool call, a single retrieval, a single formatting step. If you know OpenTelemetry, a run is the same thing as a span. Each run records its inputs, its outputs, its start and end time, any error, and a run type. The run type is one of llm, chain, tool, retriever, embedding, prompt or parser. The type matters because LangSmith shows each one differently: llm runs show token counts and cost, retriever runs show the documents returned. A function you decorate yourself defaults to chain.
A trace is all the runs that belong to one operation, shaped as a tree. When a user asks a question and your function calls a retriever and then a model, the function is the root run and the retriever and the model are its children. All of them share one trace ID. A single trace can hold at most 25,000 runs; beyond that the extras are rejected, which only ever matters for very long-running agents.
A project is the container that holds the traces of one application or service. You choose it with the LANGSMITH_PROJECT setting, and it is created automatically the first time something is sent to it. If you never set it, traces go to a project named default. One wrinkle to remember: in the API a project is called a session, so when you see session_id in a method signature it means the project's UUID, not a chat session.
A thread is a sequence of traces that form one multi-turn conversation. One turn of a chatbot is one trace, and the whole conversation is a thread. LangSmith groups traces into a thread when they carry a metadata key called session_id or thread_id with the same value. You set it yourself, and it must be on every run including the child runs.
Around these four, a few more nouns appear later in the guide. Feedback is a score attached to a run, for example "correctness: 1" or "user_thumbs: 0", with an optional comment. Tags are free-form strings on a run, and metadata is a set of key-value pairs on a run; both are searchable, and you will use them constantly. A dataset is a collection of examples, where each example has inputs and usually a reference output. An experiment is the result of running one version of your application over a dataset. An evaluator is a function or a judge model that turns an output into a score.
Finally, the organisation of your account. An organization contains one or more workspaces, and the workspaces contain your projects, prompts and datasets. Billing lives at the organization level. On a personal account you will have one organization and one workspace and can ignore all of this. It starts to matter when a key can see more than one workspace, which the configuration section covers.
- Take a simple chatbot that answers questions from a document.
- Draw the tree of runs for one question: the root, then the children, and mark each child with a run type.
- Draw how three consecutive questions would form a thread.
Creating an account and an API key
Everything that follows needs an account and a key. Sign up at https://smith.langchain.com with Google, GitHub or an email address. The Developer plan costs nothing, includes one seat and 5,000 base traces per month, and is plenty for learning. Without a payment card on the account, the free tier stops accepting traces past 5,000 in a month; adding a card raises the hourly limits and lets you go past the included traces on pay-as-you-go pricing. Pricing changes, so read the pricing page when you decide, and do not rely on any number quoted in a blog post.
Next, create an API key. In the interface open Settings, then API Keys, and create one. The key is shown once; copy it straight into a password manager or a .env file. There are two kinds, and their prefixes tell you which you hold:
| Key type | Prefix | Who makes it | Use it for |
|---|---|---|---|
| Personal access token (PAT) | lsv2_pt_ |
You | Your own scripts and notebooks. It has the same permissions you have. |
| Service key | lsv2_sk_ |
An organization admin | Applications and CI. Scoped to chosen workspaces. |
Keys never expire unless you set an expiry when creating them. Very old keys that start with ls__ were switched off on 22 October 2024; if a tutorial shows one, create a new key instead.
The last account decision is the region, and it matters more than it looks. LangSmith Cloud runs in separate regions, your organization lives in exactly one of them, and the API address differs per region. A key created in one region does not work against another region's address.
| Region | API endpoint |
|---|---|
| GCP US (the default) | https://api.smith.langchain.com |
| GCP EU | https://eu.api.smith.langchain.com |
| GCP APAC | https://apac.api.smith.langchain.com |
| AWS US | https://aws.api.smith.langchain.com |
The web interface follows the same pattern: smith.langchain.com, eu.smith.langchain.com and so on. An organization cannot be moved between regions afterwards, so choose deliberately. If you work for an employer in the Gulf or Egypt with rules about where data may live, check which region the company's data-handling policy allows before you send real prompts, since traces contain the actual text your users typed. There is no Middle East region at the time of writing, and there is no EU contracting entity, so this is a question for your legal or security team rather than a setting you can ignore. Self-hosting, which keeps everything in your own cloud account, is an Enterprise option and is covered in the Senior guide.
LANGSMITH_ENDPOINT to the regional address, with no trailing slash. A trailing slash is enough to cause the same failure.
- Sign up and open Settings, then API Keys.
- Create a personal access token named after your laptop and store it in your password manager.
- Look at the address bar and note whether your interface is on
smith.langchain.comor a regional host.
lsv2_pt_, and the region your endpoint must match.
Installing the SDK and checking the setup
The SDK is a normal library. You need Python 3.10 or newer. Always use a virtual environment, so that the install does not disturb anything else on your machine.
python3 -m venv .venv && source .venv/bin/activate
pip install -U langsmith
python -c "import langsmith; print(langsmith.__version__)"
The last line prints the installed version, which for this guide is 0.14.2 or newer. On Windows PowerShell the equivalent is:
py -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -U langsmith
python -c "import langsmith; print(langsmith.__version__)"
If you prefer uv, the equivalent of the install line is uv add langsmith. Optional extras exist for specific integrations, for example pip install -U "langsmith[pytest]" for the pytest plugin, but you do not need any of them yet.
Configuration is done with environment variables, which is a good design: nothing secret lives in your source files. Set three of them.
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY="lsv2_pt_your_key_here"
export LANGSMITH_PROJECT="my-first-app"
LANGSMITH_TRACING=true is the master switch. Without it the SDK records nothing, which is convenient for switching tracing off in tests. LANGSMITH_API_KEY is your key. LANGSMITH_PROJECT is optional, and names the project that will receive the traces. If your organization is outside the US region, also set LANGSMITH_ENDPOINT as described above.
On Windows, $env:LANGSMITH_TRACING="true" sets the variable for the current PowerShell window only. setx LANGSMITH_API_KEY "lsv2_pt_..." stores it permanently, but only new terminals see it.
You will see older material that uses LANGCHAIN_TRACING_V2, LANGCHAIN_API_KEY and LANGCHAIN_PROJECT. The SDK reads the LANGSMITH_ name first and falls back to the LANGCHAIN_ name, so the old names still work, but write the new ones in anything you create today.
.env file that is listed in .gitignore. If one leaks, delete it in Settings and create a new one.
Now check that everything is connected, with the smallest possible traced function.
from langsmith import Client, traceable
@traceable
def hello(text: str) -> str:
return text.upper()
print(hello("hi from langsmith"))
Client().flush()
Run it with python hello_trace.py. It prints HI FROM LANGSMITH, and within a few seconds a run named hello appears in the project. The last line, Client().flush(), matters: the SDK sends traces from a background thread in batches, so a short script can exit before the batch is sent. Flushing waits until it is. A long-running web server never has this problem, but every script, test and serverless function does.
If you installed the command-line tool you can also check from the terminal. The installers are below; pick the one for your system.
# macOS or Linux
curl -fsSL https://cli.langsmith.com/install.sh | sh
# macOS with Homebrew
brew install langchain-ai/tap/langsmith-cli
# then log in in the browser and list your projects
langsmith auth login
langsmith project list
On Windows use irm https://cli.langsmith.com/install.ps1 | iex in PowerShell, or scoop install langsmith-cli. If a browser is not available, for example over SSH, add --no-browser to the login command. The CLI is a convenience and everything in this guide can be done without it.
- Create a virtual environment, install
langsmithand export the three variables. - Run
hello_trace.py. - Open the LangSmith interface, choose Tracing, and open the project
my-first-app.
hello whose input is "hi from langsmith" and whose output is "HI FROM LANGSMITH".
Your first real trace: a small question-answering app
The hello example proves the plumbing. Now build something with the shape of a real application: a function that retrieves context, calls a model and returns an answer. You will need the OpenAI library and a key for it, but the same approach works for the Anthropic and Gemini clients.
pip install -U openai
export OPENAI_API_KEY="sk-..."
from openai import OpenAI
from langsmith import Client, traceable
from langsmith.wrappers import wrap_openai
# wrap_openai makes every call on this client show up as an llm run
openai_client = wrap_openai(OpenAI())
DOCS = {
"refunds": "Refunds are issued within 14 days of purchase.",
"shipping": "Standard shipping takes 3 to 5 working days.",
}
@traceable(run_type="retriever")
def get_context(question: str) -> list[dict]:
# a toy retriever: return documents whose key appears in the question
hits = [t for k, t in DOCS.items() if k in question.lower()]
return [{"page_content": t} for t in hits]
@traceable # run_type defaults to "chain"
def answer_question(question: str) -> str:
context = get_context(question)
response = openai_client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": f"Answer using only this context: {context}"},
{"role": "user", "content": question},
],
)
return response.choices[0].message.content
if __name__ == "__main__":
print(answer_question("How long do refunds take?"))
Client().flush()
Read the code for what it does, not for what it says about OpenAI. There are exactly two LangSmith ideas in it.
The first is @traceable. Put it above a function and every call to that function becomes a run, with the arguments as inputs and the return value as outputs. When one traced function calls another, the inner call is recorded as a child of the outer one automatically, because the SDK tracks the current run for you. That is how get_context ends up nested under answer_question without any wiring. The decorator works on both normal and async functions.
The second is wrap_openai. It wraps the OpenAI client so that each call to chat.completions.create is recorded as an llm run, with the messages, the response, the model name and the token counts. Equivalent wrappers exist for other providers: wrap_anthropic and wrap_gemini are in the same module. Wrapping is what gives you the cost figure, because cost is computed from the model name and the token counts the provider returned.
Run it and open the trace. You should see answer_question at the top, with get_context and the model call nested inside. Click the model call and read the actual messages sent, the answer returned, the latency, and the token counts. That panel is the thing you will look at most often in this product.
Both decorator and wrapper are opt-in and cost nothing when LANGSMITH_TRACING is unset, so you can leave them in production code and control tracing with the environment variable.
answer_question is a good name and run is a bad one. You can also set it explicitly with @traceable(name="support-answer"). A project full of well-named runs is searchable, and one full of run and wrapper is not.
- Save
qa_app.pyand run it with a question that mentions "shipping". - Open the trace and click the
llmrun. - Find the system message and confirm it contains the retrieved shipping sentence. Note the token counts and latency.
Reading a trace in the interface
Tracing is only useful if you can read the result quickly, so it is worth a section of its own. Open your project from Tracing in the sidebar. You see a table with one row per trace, and columns for the name, the input, the output, the start time, the latency, the token count and the cost. Click a row and a panel opens with the run tree on the left and the details of the selected run on the right.
Start with the tree. Each node is a run, indented under its parent, with its duration next to it. The slowest node is usually obvious, and the waterfall view shows the durations as bars against a timeline so you can see whether steps ran one after another or in parallel. A red marker means the run raised an error, and errors propagate upward, so if a child failed you will see the parent marked too. The deepest red node is normally the real culprit.
Then read the details of one run at a time. Input and Output are the arguments and the return value. For an llm run they appear as chat messages, so you read the conversation as the model saw it. Metadata shows the model name, the provider and anything you added yourself. Feedback shows scores attached to the run. The token and cost figures sit in the header.
The table is also a search box. The filter bar accepts conditions on name, status, latency, tags, metadata and more. The same condition language is available from the SDK and the CLI, written as functions, for example and(eq(run_type, "llm"), gt(latency, 5)), meaning "model calls slower than five seconds". You will meet it again in the commands section.
Three habits make the interface much faster. Filter before you scroll: filtering to errors or slow runs turns a thousand rows into ten. Open the failing run's parent as well as the run itself, since the cause is often in what the parent sent. And copy the trace link to share it: a colleague with access to the workspace can open the same tree, which is worth more than any screenshot. The link to a run comes from the interface; do not build it by hand.
- Call
answer_questionfive times with different questions, including one that matches nothing inDOCS. - In the project table, sort by latency and open the slowest trace.
- Use the filter bar to show only runs whose name is
get_context.
Tracing options: names, tags, metadata and run types
Once traces arrive, the next question is how to make them findable. A project of a thousand anonymous runs is a haystack; the same thousand runs with good tags and metadata is a database you can query. The decorator takes everything you need.
from langsmith import traceable
@traceable(
name="support-answer",
run_type="chain",
tags=["support", "v2"],
metadata={"app_version": "2.1.0", "customer_tier": "free"},
)
def answer(question: str) -> str:
return f"You asked: {question}"
# values that only exist at call time go in langsmith_extra
answer(
"where is my order?",
langsmith_extra={"metadata": {"user_id": "u-1042"}, "tags": ["beta"]},
)
The two words to keep apart are tags and metadata. A tag is a label with no value: support, v2, beta. Use tags for categories you filter by. Metadata is a key with a value: app_version: 2.1.0, user_id: u-1042. Use metadata for facts you want to group or search by, such as a version, a customer tier or a model setting. Values known when you write the code go in the decorator; values known only at call time, such as the user's ID, go in langsmith_extra, which the decorator strips out before calling your function, so your function signature does not change.
The decorator accepts several more keyword arguments. client lets you supply a specific Client object. project_name sends this function's traces to a different project than the default, which is handy in a monorepo. process_inputs and process_outputs let you rewrite the payload before it leaves your machine, which the configuration section uses for redaction. enabled switches tracing on or off for one function. reduce_fn tells the SDK how to combine the chunks of a streamed output into a single stored output.
Choosing a run type is worth a moment. If you leave it, a function is a chain. If your function calls a model through a library that is not wrapped, set run_type="llm" so that the interface treats it as a model call. For a custom model call to show cost, put the provider and model name in metadata as ls_provider and ls_model_name and report token counts in usage_metadata. For a retrieval function, run_type="retriever" makes the returned documents display as documents. For a function that calls an external API, run_type="tool". The type is only a label, but it decides how the run is drawn and what is computed for it.
If you need to trace a block of code instead of a whole function, use the trace context manager, which records a run for the code inside the with block:
import langsmith as ls
def handle(question: str) -> str:
with ls.trace(name="post-process", run_type="chain", inputs={"q": question}) as run:
result = question.strip().lower()
run.end(outputs={"result": result})
return result
To switch tracing off for a region of code, even when the environment variable says it is on, use tracing_context:
import langsmith as ls
with ls.tracing_context(enabled=False):
# nothing inside this block is sent to LangSmith
print("private work")
The precedence, from strongest to weakest, is the tracing_context block, then a setting on the client or the decorator, then the environment variable. That ordering is why you can leave LANGSMITH_TRACING=true set for a whole service and still carve out an exception.
app_version, environment, user_id. Six months from now the question will be "did the error rate change after version 2.1?", and that question is only answerable if every trace carries the version.
- Add
tagsandmetadatatoanswer_questioninqa_app.py, including alangsmith_extrauser ID at call time. - Run it three times with different user IDs.
- In the interface, filter the project by one of the user IDs.
Tracing from LangChain, LangGraph and TypeScript
You can reach LangSmith from several places besides a decorated Python function. The right choice depends on what you already use.
If your application is written with LangChain or LangGraph, you do not change any code. Set the same three environment variables and every chain, model call, tool call and graph node is traced automatically, nested in the order they ran. This is the reason many people meet LangSmith first: the integration is invisible. The LangGraph guide shows what the graphs look like, and each node of a graph appears as a run in the trace.
If you use TypeScript or JavaScript, install the package and use the same environment variables:
npm install langsmith openai
export LANGSMITH_TRACING=true
export LANGSMITH_API_KEY="lsv2_pt_..."
export LANGSMITH_PROJECT="my-first-app"
import { traceable } from "langsmith/traceable";
import { wrapOpenAI } from "langsmith/wrappers";
import OpenAI from "openai";
const openai = wrapOpenAI(new OpenAI());
const answerQuestion = traceable(
async (question: string) => {
const res = await openai.chat.completions.create({
model: "gpt-4o-mini",
messages: [{ role: "user", content: question }],
});
return res.choices[0].message.content;
},
{ name: "answer_question", tags: ["support"] }
);
console.log(await answerQuestion("How long do refunds take?"));
In TypeScript, traceable is a function that wraps your function rather than a decorator, and the options object comes second. In a serverless environment such as a function that returns as soon as it has a response, the batch may not be sent before the runtime freezes. Set LANGSMITH_TRACING_BACKGROUND=false so that traces are sent inline, or call await client.awaitPendingTraceBatches() on a Client before returning.
If you use a different framework, check the integrations list in the docs first. There are integrations for the OpenAI Agents SDK, the Vercel AI SDK, CrewAI, Pydantic AI, Google ADK, LiteLLM and more, plus OpenTelemetry for anything else that can emit spans. Reaching for a documented integration is nearly always less work than wrapping things by hand.
Use the integration when
- You already use LangChain or LangGraph: set the variables and you are done
- The provider client has a wrapper:
wrap_openai,wrap_anthropic - The framework is in the integrations list
Write manual tracing when
- You have your own functions that matter: retrievers, tools, business rules
- You call a model through a raw HTTP request
- You need a run that does not correspond to one function
Most real applications mix the two: the framework or wrapper traces the model calls, and @traceable covers your own code so the tree tells the whole story.
- If you have Node.js installed, create a folder, run
npm install langsmith openaiand save the TypeScript example asqa.ts. - Run it with a TypeScript runner such as
npx tsx qa.ts. - Open the project and compare the TypeScript trace with the Python one.
Threads: following a conversation
A chatbot trace shows one turn. The bug you are chasing is usually a whole conversation: the assistant forgot what the user said three messages ago, or the second answer contradicts the first. To see the conversation, LangSmith needs to know which traces belong together, and you tell it with a thread identifier.
The rule is simple. Give every trace in one conversation the same value under the metadata key session_id, thread_id or conversation_id, and put it on the runs themselves. The docs recommend a UUID, version 7 being the preferred form. Keep it stable for the whole conversation and generate a new one for the next.
import uuid
from langsmith import Client, traceable
thread_id = str(uuid.uuid4()) # one per conversation
@traceable(run_type="chain", name="chat-turn")
def chat_turn(message: str, history: list[dict]) -> str:
# call your model here; this echo keeps the example runnable
return f"You said: {message}"
history: list[dict] = []
for text in ["Hello", "What is your refund policy?", "Thanks"]:
reply = chat_turn(
text,
history,
langsmith_extra={"metadata": {"thread_id": thread_id}},
)
history += [{"role": "user", "content": text}, {"role": "assistant", "content": reply}]
Client().flush()
After it runs, open the Threads tab in your project. The three traces appear as one thread, and opening it shows the whole exchange as a conversation. The key has to be on the runs LangSmith looks at, so if you trace child runs yourself, pass the same metadata to those as well.
Threads are what make conversation-level questions possible: how many turns did users need before giving up, in which thread did the model start repeating itself. The Mid guide builds evaluators that work on whole threads. At this level, the skill to learn is simply to set the identifier from the first turn, because you cannot add it retroactively to traces that have already been stored without it.
One product note that catches people using old tutorials: the Turns view in the thread detail page is deprecated and is scheduled to be removed from Cloud on 31 October 2026. Use the Messages or Details view instead.
- Run
chat_thread.py. - Open the project's Threads tab and open the thread.
- Change the code so the thread ID is generated inside the loop, run again, and look at the Threads tab.
Feedback: attaching scores to runs
A trace tells you what happened. Feedback tells you whether it was good. A feedback item is attached to a specific run and has a key (the name of the quality, such as correctness or user_score), a score (a number) or a value (a category), and an optional comment. The source can be a person clicking in the interface, a thumbs-up button in your own product, an automatic evaluator, or a reviewer in an annotation queue.
From code, you attach feedback to a run by its ID. For that you need the run ID, which you can set yourself before calling the function, through langsmith_extra:
import uuid
from langsmith import Client, traceable
client = Client()
@traceable(name="support-answer")
def answer(question: str) -> str:
return "Refunds take 14 days."
run_id = uuid.uuid4()
answer("How long do refunds take?", langsmith_extra={"run_id": run_id})
client.flush()
# later, when the user clicks thumbs-up in your product
project = client.read_project(project_name="my-first-app")
client.create_feedback(
run_id,
key="user_score",
score=1.0,
comment="Helpful",
session_id=project.id,
)
Note session_id, which is the project's UUID, found by reading the project by name. Without it the SDK warns that creating feedback for a run without a session ID is deprecated and will stop working, and on the newer datastore that LangSmith is migrating to it raises a ValueError saying the deployment cannot locate the run. Passing it is cheap, so pass it.
The feedback then appears on the run in the interface, and you can filter and chart by it. A typical use is to record a user's thumbs-up or thumbs-down and then open only the thumbs-down traces to find what went wrong.
extend_trace_retention=True upgrades the whole trace to extended retention, 180 days, at a higher price. That is useful when you deliberately want to keep the traces users complained about. Do not set it on every call.
You can also give feedback directly in the interface: open a run, use the feedback controls, and enter a score. That is how you start building a sense of quality before you have any automation.
- Run
feedback_example.py, after editing the project name to match yours. - Open the trace for the
support-answerrun. - Find the feedback entry named
user_score, then add a second one by hand in the interface.
Finding traces from code and from the command line
Sooner or later you will want traces outside the browser: to count errors, to pull the last ten failures into a notebook, or to give an AI coding assistant access. There are two routes, the SDK and the CLI.
There is a change here that older tutorials do not mention. LangSmith introduced a new trace datastore called SmithDB in 2026, and the older query methods are deprecated in favour of new ones. The deprecated client.list_runs() and client.read_run() still exist in SDK 0.14.2 but emit a warning, "<method>() is deprecated and will be removed after Jan 31, 2027.", and they stop working when the Cloud backend removes them. The new methods live in client.runs, client.traces and client.threads, and need Python langsmith 0.10.15 or newer.
from datetime import datetime, timedelta, timezone
from langsmith import Client
client = Client()
project = client.read_project(project_name="my-first-app")
for run in client.runs.query(
project_ids=[project.id],
min_start_time=datetime.now(timezone.utc) - timedelta(days=7),
selects=["NAME", "STATUS", "INPUTS"],
filter='eq(run_type, "llm")',
page_size=20,
):
print(run)
Four details decide whether this works. project_ids takes project UUIDs; the old project_name argument was removed, which is why the example reads the project first. min_start_time defaults to one day ago, so a query without it silently covers only the last 24 hours and looks as if your older traces are missing. selects lists the fields you want, written in uppercase, and the default is to return only the id, so forgetting it gives you a list of bare IDs. And the run_type value in a filter is written in the filter language shown above. Whether the loop is sync or async depends on the client; in async code use async for run in client.runs.query(...).
The command-line tool is often quicker for looking around:
langsmith project list
langsmith trace list --project my-first-app --limit 5
langsmith trace get <trace-id> --project my-first-app --full
langsmith run list --project my-first-app --run-type llm
langsmith thread list --project my-first-app
langsmith --format json trace list --project my-first-app --limit 5 -o traces.json
trace list defaults to the last seven days and 20 rows, and run list defaults to 50 rows, so a small number of results is not necessarily a problem. thread list requires --project. The --format json flag and the -o file option make the output easy to hand to other tools. Useful filters include --error to show only failures, --min-latency for slow runs and --tags to select by tag.
If a CLI command fails with an authentication error, run langsmith auth login again, or make sure LANGSMITH_API_KEY is exported. Note that an exported LANGSMITH_ENDPOINT overrides the address stored in a CLI profile, which can silently send a command to the wrong place.
- Install the CLI and run
langsmith project list. - Run
langsmith trace list --project my-first-app --limit 5. - Run
query_runs.pyand then removeselectsfrom it and run it again.
Datasets and your first evaluation
Tracing shows what your application did. Evaluation answers a different question: is this version better than the last one? You cannot answer it by eyeballing a few traces, because improving one answer often breaks another. The fix is the same as in ordinary software: a fixed set of test cases that you run every time.
A dataset is that set. It is a collection of examples, and each example has inputs, which are a dictionary your application accepts, an optional reference outputs dictionary holding the answer you expect, and optional metadata. An experiment is what you get when you run your application over a dataset: the outputs, the scores and the traces for every example, all saved and comparable. An evaluator is the function that produces the scores.
Here is the smallest complete evaluation. The target is a plain function, so this runs without any model key and you can see the machinery clearly.
from langsmith import Client
client = Client()
dataset = client.create_dataset(
dataset_name="refund-faq",
description="Questions about refunds and shipping",
)
client.create_examples(
dataset_id=dataset.id,
examples=[
{"inputs": {"question": "How long do refunds take?"},
"outputs": {"answer": "14 days"}},
{"inputs": {"question": "How long does shipping take?"},
"outputs": {"answer": "3 to 5 working days"}},
],
)
def target(inputs: dict) -> dict:
# replace this with a call to your real application
q = inputs["question"].lower()
if "refund" in q:
return {"answer": "Refunds are issued within 14 days."}
return {"answer": "Standard shipping takes 3 to 5 working days."}
def contains_reference(outputs: dict, reference_outputs: dict) -> dict:
ok = reference_outputs["answer"].lower() in outputs["answer"].lower()
return {"key": "contains_reference", "score": 1 if ok else 0}
results = client.evaluate(
target,
data="refund-faq",
evaluators=[contains_reference],
experiment_prefix="first-eval",
max_concurrency=2,
)
Run it once. If you run it a second time, create_dataset will complain that the dataset already exists, so for repeated runs delete the dataset in the interface or comment out the creation lines.
There are four ideas to take from it. The target function takes the example's inputs dictionary and returns a dictionary of outputs; LangSmith calls it once per example. An evaluator is an ordinary function whose parameters are chosen by name: it may ask for inputs, outputs, reference_outputs, run or example, and LangSmith passes what you ask for. It returns a boolean, a number, or a dictionary with a key and a score (or a categorical value) and an optional comment. experiment_prefix gives the experiment a readable name, and LangSmith adds a unique suffix. And max_concurrency controls how many examples run at once; keep it low until you know your model provider's rate limits.
Open Datasets & Experiments, choose refund-faq, and open the experiment. You see one row per example with the output, the contains_reference score and a link to the full trace of that example. Every evaluation run is itself traced, so a low score is always one click from the exact steps that produced it.
Checking whether a string contains another is a weak evaluator for free-form answers. A more capable kind is an LLM-as-judge: a second model reads the question, the answer and the reference and returns a score. The separate openevals package provides ready-made ones such as create_llm_as_judge, and the Mid guide covers them and their pitfalls. For now, remember the taxonomy: human, code and LLM-judge evaluators exist; offline evaluation runs on a dataset before you deploy; online evaluation scores real production traces as they arrive, without reference answers.
- Run
first_eval.pyand open the experiment in the interface. - Change
targetso the refund answer says "30 days" and delete the dataset's old experiment by re-running after dropping the creation lines. - Compare the two experiments in the dataset's experiment list.
Prompts and the Playground
Prompt text is the part of an LLM application that changes most often, and the part least suited to living inside source code, because a wording tweak then requires a code release. LangSmith lets you keep prompts as versioned objects and pull them at run time.
A prompt is a template with variables, either a chat template with system and user messages or a plain completion template. Every time you save a change it creates a commit, an immutable version identified by a hash. A commit tag is a movable label on a commit; staging and production are reserved names used for environments, so you can promote a version by moving the tag rather than changing code.
To create and use a prompt from code you need the langchain-core package for the template class, and the langsmith client to push and pull.
pip install -U langchain-core
from langchain_core.prompts import ChatPromptTemplate
from langsmith import Client
client = Client()
prompt = ChatPromptTemplate.from_messages([
("system", "You are a support agent. Answer in one sentence."),
("user", "{question}"),
])
# creates the prompt the first time, a new commit afterwards
url = client.push_prompt("support-answer-prompt", object=prompt)
print(url)
latest = client.pull_prompt("support-answer-prompt")
print(latest.invoke({"question": "How long do refunds take?"}))
You address a prompt by name, optionally followed by a colon and either a commit hash or a tag: support-answer-prompt, support-answer-prompt:12344e88, support-answer-prompt:production. A prompt you do not own and that has been shared publicly is addressed as owner/name. Pulling the tag production is how an application gets "the approved version" while a colleague experiments with newer commits in the interface.
Pulled prompts are cached in your process, on by default, with a five-minute time to live and a background refresh every minute. That is usually what you want, but it explains a puzzle for beginners: you edit a prompt in the interface and your running program seems to ignore it for a few minutes. Wait, restart, or pass skip_cache=True to pull_prompt while testing.
The Playground is the interface for trying a prompt without writing code. You open a prompt, choose a model, fill in variables, run it, and change the text and run again. It can also run your prompt over a dataset, which gives you an evaluation without a script. You supply the model provider's API key in the workspace settings as a secret, so that calls are made on your behalf. The Playground is the best place for a non-engineer, such as a product manager or a domain expert, to improve wording while you keep the evaluation rigour.
production tag. The application code never changes and the rollback is moving a tag back.
- Run
prompts.pyand open the printed URL. - In the interface, edit the system message and save, creating a second commit.
- Pull
support-answer-promptagain, and then pull the first commit's hash explicitly.
Configuration and privacy: what leaves your machine
Every trace carries the real inputs and outputs of your program. For a chatbot that means the user's messages; for a healthcare or HR assistant it may mean personal data. Decide what you send before you send it, not after.
The environment variables you will use most:
| Variable | What it does |
|---|---|
LANGSMITH_TRACING |
true turns tracing on. Leave it unset in tests. |
LANGSMITH_API_KEY |
Your PAT or service key. |
LANGSMITH_PROJECT |
The target project. Defaults to default. |
LANGSMITH_ENDPOINT |
The API address. Required outside the US region. |
LANGSMITH_WORKSPACE_ID |
Needed when the key can see more than one workspace. |
LANGSMITH_TRACING_SAMPLING_RATE |
A number from 0 to 1 that keeps only that share of traces. |
LANGSMITH_HIDE_INPUTS / LANGSMITH_HIDE_OUTPUTS |
true removes the whole payload. |
The simplest privacy control is hiding everything: with LANGSMITH_HIDE_INPUTS=true and LANGSMITH_HIDE_OUTPUTS=true, the structure, timings and errors are still recorded but the content is not. That is a good default while you work out a better policy, though it makes traces far less useful for debugging. The better middle path is redacting only sensitive fields, with a function applied per traced function:
import re
from langsmith import traceable
def scrub(inputs: dict) -> dict:
text = inputs.get("message", "")
text = re.sub(r"[\w.+-]+@[\w-]+\.[\w.]+", "[EMAIL]", text)
return {**inputs, "message": text}
@traceable(process_inputs=scrub)
def handle(message: str) -> str:
return "ok"
Two behaviours are worth knowing. A function-level process_inputs or process_outputs takes precedence over client-level settings. And since SDK 0.12 these redactors fail closed: if your redactor raises an error, the payload is replaced by a message such as Processing failed; inputs dropped (process_inputs: KeyError) instead of sending the unredacted data. That is the safe behaviour, and a trace showing it means your redaction function has a bug to fix, not a setting to disable.
For volume, sampling keeps only a fraction of traces. Setting LANGSMITH_TRACING_SAMPLING_RATE=0.25 records about a quarter of requests. Since SDK 0.12 sampling is deterministic. It is a cost lever for busy production services and unnecessary while you are learning.
If your key spans several workspaces, for instance an organization-scoped service key, LangSmith cannot guess which workspace you mean. Set LANGSMITH_WORKSPACE_ID to the workspace's UUID, or the requests fail with a 403.
For notebooks there is one more trap. The Python SDK caches environment variable reads for speed. If you change LANGSMITH_PROJECT or the key in a running Jupyter kernel and then see traces in the wrong place, clear the cache and reload your .env:
from dotenv import load_dotenv
from langsmith import utils
utils.get_env_var.cache_clear()
load_dotenv(override=True)
- Apply
scrubto a function in your app and call it with a message containing an email address. - Open the trace and read the input.
- Set
LANGSMITH_HIDE_INPUTS=true, run again and compare.
[EMAIL] in place of the address, then an input that is hidden altogether. You now control what is stored.
Common errors and how to read them
Most LangSmith failures are silent by design. Tracing runs in a background thread and is built never to crash your application, so when something is wrong you usually see a warning in the logs and no traces in the interface. Learn the handful of messages below and you will diagnose nearly everything.
LangSmithMissingAPIKeyWarning: API key must be provided when using hosted LangSmith API. The key is not set in the process that is running. Check echo $LANGSMITH_API_KEY in the same terminal, and in a notebook clear the environment cache as shown earlier.
LangSmithAuthError: Authentication failed for /runs/multipart, or a log line Failed to multipart ingest runs. An HTTP 401. The causes, in order of likelihood: the endpoint belongs to a different region than your organization, there is a trailing slash on LANGSMITH_ENDPOINT, the key was deleted or deactivated, or it is an old ls__ key. Deactivation takes up to a minute to apply because authentication is cached.
LangSmithUserError: This API key is org-scoped and requires workspace specification. You used an organization-wide service key. Set LANGSMITH_WORKSPACE_ID or pass Client(workspace_id=...).
LangSmithRateLimitError: Rate limit exceeded. An HTTP 429. The SDK retries with back-off, so isolated ones are harmless. Persistent ones mean you hit an hourly event or data limit, or the free tier's monthly cap of 5,000 traces. Sample your traces or raise the plan.
LangSmithConnectionError: Connection error caused failure to POST /runs. The SDK could not reach the server. Check the endpoint, your proxy and your network. If your employer inspects HTTPS traffic with its own certificate, trust that certificate authority in your environment instead of turning verification off. A variant says the content length exceeds the maximum size limit, which means a single payload is too large; send less data per run.
The traces never show up, but no error appears. The usual reason is that the process exited before the background batch was sent. Call Client().flush() at the end of scripts and tests. In LangChain code, call wait_for_all_tracers() from langchain_core.tracers.langchain. In serverless TypeScript, set LANGSMITH_TRACING_BACKGROUND=false.
Traces appear, but in the wrong project. Either LANGSMITH_PROJECT was not set where you think it was, or the notebook cache explains it. A project_name on a decorator also overrides the environment variable for that function.
A query returns only one day of data, or only IDs. Those are the defaults of client.runs.query, as described earlier.
A deprecation warning saying a method "will be removed after Jan 31, 2027". You used a legacy query method. It works for now. Plan to move to the runs, traces and threads methods.
SSL_CERT_FILE or REQUESTS_CA_BUNDLE, or to ask your IT team for it.
- Set
LANGSMITH_ENDPOINTtohttps://api.smith.langchain.com/with a trailing slash, or to a wrong region, and runhello_trace.py. - Read the log output and find the authentication message.
- Fix the endpoint and run again.
Retention, limits and cost: the parts that surprise people
A tracing tool that costs nothing on day one can cost something by month three, so understand the levers early.
Traces have a retention period. On LangSmith Cloud, base traces are kept for 14 days, and extended traces for 180 days. The extended period was 400 days until 14 September 2026, when it was reduced for new traces. A trace is upgraded from base to extended in three situations: an online evaluator scores it, an automation rule that extends retention matches it, or your code sends feedback with extend_trace_retention=True. Upgrades cost more, and on new evaluators and rules the option is switched on by default, so read those settings when you create them. Experiments are created with extended retention.
Datasets are not affected by any of this: they are kept indefinitely. So the practice is to treat traces as temporary and promote the ones you care about into a dataset.
The Developer plan includes 5,000 base traces a month; Plus costs $39 per seat per month with 10,000 included. Beyond that you pay per trace. The documentation says to read per-trace prices from the pricing page, and this guide does not repeat them because they change. There are also hourly limits on the number of events and the volume of data, which differ per plan. You can set monthly usage limits per workspace, project or user, so a bug that loops cannot produce a surprise invoice.
Two technical limits matter to beginners. One trace can contain at most 25,000 runs. And the ingestion API has per-minute rate limits per key, which the SDK respects by batching up to 100 runs per call.
For a new project, three habits keep cost and risk low: use a separate project per application, so that you can see and delete traces by application; sample production traffic once it is high enough that storing everything is wasteful; and do not send data you would not want stored. The Senior guide treats cost governance as a platform topic.
- Open the billing or usage page in your organization's settings.
- Find the number of base traces used this month and the monthly limit.
- Find where a workspace usage limit could be set.
Putting it all together: a small support bot, traced and tested
Now combine the pieces in one project you could show a colleague. The goal is a tiny support assistant that answers from a few documents, traces every step with tags and metadata, groups a conversation into a thread, records user feedback, and has a dataset and an evaluation that run before each change.
Set up the project layout:
mkdir support-bot && cd support-bot
python3 -m venv .venv && source .venv/bin/activate
pip install -U langsmith openai python-dotenv
printf 'LANGSMITH_TRACING=true\nLANGSMITH_PROJECT=support-bot\n' > .env
echo '.env' > .gitignore
Add your keys to .env by hand, so they never enter the shell history: LANGSMITH_API_KEY and OPENAI_API_KEY. Then the application:
import uuid
from dotenv import load_dotenv
load_dotenv()
from openai import OpenAI
from langsmith import Client, traceable
from langsmith.wrappers import wrap_openai
llm = wrap_openai(OpenAI())
ls_client = Client()
DOCS = {
"refund": "Refunds are issued within 14 days of purchase.",
"shipping": "Standard shipping takes 3 to 5 working days.",
"warranty": "All devices have a one-year warranty.",
}
@traceable(run_type="retriever", name="lookup-docs")
def lookup(question: str) -> list[str]:
return [t for k, t in DOCS.items() if k in question.lower()]
@traceable(name="support-answer", tags=["support"], metadata={"app_version": "1.0"})
def answer(inputs: dict) -> dict:
question = inputs["question"]
context = lookup(question)
res = llm.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system",
"content": "Answer in one sentence using only this context: " + " ".join(context)},
{"role": "user", "content": question},
],
)
return {"answer": res.choices[0].message.content}
def chat(questions: list[str]) -> None:
thread_id = str(uuid.uuid4())
for q in questions:
out = answer(
{"question": q},
langsmith_extra={"metadata": {"thread_id": thread_id}},
)
print(q, "->", out["answer"])
ls_client.flush()
if __name__ == "__main__":
chat(["How long do refunds take?", "And the warranty?"])
And a separate script that turns the same function into a tested one:
from dotenv import load_dotenv
load_dotenv()
from langsmith import Client
from bot import answer
client = Client()
NAME = "support-bot-eval"
if not client.has_dataset(dataset_name=NAME):
ds = client.create_dataset(dataset_name=NAME)
client.create_examples(
dataset_id=ds.id,
examples=[
{"inputs": {"question": "How long do refunds take?"},
"outputs": {"answer": "14 days"}},
{"inputs": {"question": "How long is the warranty?"},
"outputs": {"answer": "one year"}},
],
)
def mentions_reference(outputs: dict, reference_outputs: dict) -> dict:
hit = reference_outputs["answer"].lower() in outputs["answer"].lower()
return {"key": "mentions_reference", "score": int(hit)}
client.evaluate(
answer,
data=NAME,
evaluators=[mentions_reference],
experiment_prefix="v1",
max_concurrency=2,
)
The workflow is then a loop you will repeat for the rest of your career with this tool. Run python bot.py and read the traces: does the retriever return the right documents, does the prompt contain them, is the answer faithful? Fix what you find, for example by making lookup smarter. Run python evaluate_bot.py and compare the new experiment with the previous one in the interface. When a real user reports a bad answer, open that trace, add it to the dataset, and the bug has become a permanent test. Change the app version in the metadata with each release, so that every trace records which code produced it.
- Build the project above and run
bot.py, thenevaluate_bot.py. - Break the retriever on purpose by renaming a key in
DOCS, and run the evaluation again. - Compare the two experiments, then open a failing example's trace.
What you can now do, and what comes next
You now hold the working vocabulary and the core loop of LangSmith. You can create an account and a key, choose the right regional endpoint, install the SDK, and send traces from Python or TypeScript using @traceable, traceable, the provider wrappers or a framework integration. You can read a trace, find the slow or failing step, filter with tags and metadata, group a conversation into a thread, and attach feedback. You can query traces from the SDK and the CLI, taking care of the defaults of the new query method. You can build a dataset, run an experiment and read the scores, version a prompt, and control what data leaves your machine. You know the error messages that matter and the limits that cost money.
Three things to carry forward. Tracing is cheap to add and expensive to lack, so instrument from the first day, not after the first incident. Metadata is what makes traces useful later, so decide your keys early. And evaluation is a loop, not an event: production failures feed the dataset, and the dataset guards the next release.
Here is what the next levels add. The Mid guide covers writing good evaluators including LLM judges, running evaluation in pytest and CI, online evaluation and automation rules, annotation queues, feedback from end users at scale, cost tracking for custom models, prompt caching, and tracing across services with OpenTelemetry and distributed tracing. The Senior guide covers running LangSmith as a platform for several teams: workspaces and access control, the SmithDB migration at scale, self-hosting on Kubernetes with Helm, security and data residency, cost governance and upgrade policy. Meanwhile, compare what you learned here with the Langfuse guide and the evaluation-metrics approach in the RAGAS guide, which together show which ideas are universal across LLM observability and evaluation tools.
Sources
- LangSmith documentation home
- Observability quickstart
- Observability concepts
- Create an account and API key
- Regions FAQ
- Threads
- Conditional tracing
- Sample traces
- Mask inputs and outputs
- Serverless environments
- Troubleshooting variable caching
- Evaluation concepts
- Evaluation quickstart
- Evaluate an LLM application
- Prompt engineering concepts
- Manage prompts programmatically
- LangSmith CLI
- SmithDB SDK migration
- Usage and billing
- LangSmith changelog
- Python SDK on PyPI
- JavaScript SDK on npm
- LangSmith SDK releases