This is part one of three. It covers everything you need to start doing real work with Opik, not a teaser. By the end you will have installed the SDK, pointed it at either Opik Cloud or a server on your own laptop, traced a small LLM application so that every step shows up in a web page, attached tags and scores to what you traced, built a dataset of test questions, and run an experiment that scores your application against it. Mid-level and Senior take the same topics further; nothing here is thrown away.
This guide is written against Opik 2.2.86, released on 30 September 2026. Opik ships several patch releases a week, and the server, the Python SDK, the TypeScript SDK and the Helm chart all share one version number. So treat "2.2.x" as the version you are learning and run pip install -U opik whenever you start a new project. Where something changed recently and old tutorials will mislead you, this guide says so.
Each section ends with a Try it task. Do them as you go. They take a few minutes each, and these ideas only stick once you have watched your own trace appear in the interface and then broken it on purpose.
What Opik is, and the problem it solves
Opik is an open-source platform (Apache-2.0 licence) from Comet for watching, testing and improving applications built on large language models. An LLM application is any program that sends text to a model such as GPT, Claude or Gemini and does something with the answer: a chatbot, a document summariser, a retrieval-augmented question-answering service, an agent that calls tools. Opik records what those programs do, lets you score how well they did it, and keeps the prompts and test data you use to make them better.
The reason such a tool exists is that LLM applications fail in a way ordinary software does not. A normal function either returns the right number or it does not, and a unit test can tell you which. An LLM application returns fluent, confident text every time, and whether that text is correct, relevant, safe and grounded in your documents is a judgement call. A program can run without a single exception for a month while quietly telling customers something false. Nothing crashes, so nothing alerts you.
Opik attacks that problem from three directions, and these three directions are the structure of the rest of this guide.
Observability means being able to see what happened. Every request through your application becomes a record you can open: the question, the prompt that was actually sent, the documents retrieved, the model's answer, how many tokens it used, what it cost, and how long each step took. When a user says "the bot gave me a weird answer yesterday", you search for it and read the exact chain of events rather than guessing.
Evaluation means being able to measure whether the application is any good. You collect a set of example questions, run your application over all of them, and score the answers automatically. When you change the prompt or swap the model, you run the same set again and compare numbers instead of trusting a feeling.
Improvement and protection cover the rest: a library where prompts are versioned like code, annotation tools so a human can review real answers, online scoring that checks production traffic as it arrives, alerts that message your team when error rates or costs spike, and optimisers that propose better prompts. Some of these (guardrails that block unsafe output inline, for instance) are enterprise features, and this guide only mentions them so that you know they exist.
Opik is available in three forms, and it is worth understanding the difference early because it changes which steps you follow. Opik Cloud is hosted by Comet, has a free tier, and needs nothing installed except the SDK. The self-hosted version is the same open-source code running on your own machine or cluster; the local Docker Compose setup is ideal for learning, and Kubernetes with Helm is the production route. Comet Enterprise adds single sign-on, fine-grained roles and similar controls. The SDK you write code against is identical in all three. Only the address it talks to changes.
For readers in the Gulf and Egypt, the choice between Cloud and self-hosted is often a data-residency question rather than a technical one. If your employer cannot send customer conversations to a vendor's region, a self-hosted Opik inside your own cloud account or data centre keeps prompts and outputs within your boundary. You will see where that choice is made in one configuration line.
- Think of an LLM feature you have used or built: a chatbot, a summary button, a search box with an AI answer.
- Write down three ways it could give a wrong answer without raising an error.
- For each, write what you would need to see in order to find out why.
What people did before
Before dedicated tooling, most teams handled LLM applications with print statements, log files and a spreadsheet. The habit is understandable, and it works for the first afternoon. It stops working when the application grows, and seeing where it stops explains why Opik is designed the way it is.
A print of the model's answer tells you the final text but not the intermediate steps. Modern applications are rarely a single call. A retrieval-augmented generation (RAG) service embeds the question, searches a vector database, builds a prompt from the top results, calls the model, and post-processes the reply. An agent may loop through several model calls and tool calls before answering. When the final answer is wrong, the fault could sit in any of those steps, and a flat log does not show which step fed which.
Ordinary application-performance monitoring tools record timing and errors well, but they do not understand tokens, prompts, model names or the price of a call, and they have no concept of "this answer was 0.3 on faithfulness". Teams bolted those on by hand, differently every time.
The spreadsheet of test questions is the other half. People kept a column of inputs, pasted in the outputs by hand after each change, and eyeballed the differences. It is slow, it does not scale past twenty rows, and nobody can reproduce last week's run because nobody saved which prompt and model produced it.
Opik replaces the flat log with a tree of recorded steps and the spreadsheet with datasets and experiments that save the inputs, the outputs, the scores and the configuration of every run. Neighbouring tools solve parts of the same problem. If you later meet Langfuse, it is another open-source observability and evaluation platform with the same basic ideas (see the Langfuse guide), and RAGAS is a library of evaluation metrics for retrieval-augmented systems (see the RAGAS guide). Opik covers tracing, datasets, metrics, prompts and monitoring in one place, which is why it makes a good first tool to learn.
- Open the official documentation home page,
https://www.comet.com/docs/opik/, and scan the left-hand navigation. - Find the headings for Tracing, Evaluation, Prompt management and Production monitoring.
The mental model: workspace, project, trace, span
Opik has a handful of nouns, and they nest. Learn the nesting once and the interface stops being a maze.
tenant boundary: members, keys, limits
one agent or app
one interaction
steps, as a tree
A workspace is the outermost boundary you will deal with. On Opik Cloud it holds your members, your AI-provider keys and your rate limits. On a self-hosted open-source install there is a single workspace called default and no user management at all, which matters for security and is covered in the Senior level.
A project represents one agent or application. In Opik 2.x, a project scopes not only the records of what your application did but also the datasets, experiments, prompts, online scoring rules, alerts and dashboards that belong to it. If you run a customer-support bot and a document summariser, you create two projects, and each keeps its own tests and prompts. If you do not choose a project, everything lands in one literally named Default Project.
That project-scoping is the biggest change in Opik 2.0 (released in spring 2026). In the 1.x series, only traces lived in projects, while datasets, prompts and experiments were shared across the whole workspace. Old blog posts and videos describe the 1.x behaviour. If an example you find online creates a dataset without a project_name and you cannot see it where you expect, this is why.
A trace is the complete record of one interaction with your application, one request and its response. It has a unique identifier, an input, an output, start and end times, optional metadata and tags, token usage and cost, and feedback scores. The guidance from the documentation is clear and worth memorising: a trace is one user interaction or one business operation. It is not a single function call, and it is not an entire session.
A span is one operation inside a trace. A trace for a RAG request might contain a span for retrieving documents, a span for building the prompt and a span for calling the model. Spans nest: a span can contain child spans, so the whole trace is a tree. Each span has a type. The common values are general (any ordinary step), llm (a call to a model), tool (a tool or function an agent uses) and guardrail. The type matters because only spans of type llm that carry a model name, a provider and token usage get a cost computed.
Here is what a trace for a tiny question-answering app looks like conceptually:
Trace: answer_question input: "What is our refund window?" output: "30 days."
├── Span: retrieve_docs type: general (searches the knowledge base)
├── Span: build_prompt type: general (assembles instructions + context)
└── Span: call_model type: llm model: gpt-4o-mini tokens: 412 cost: $0.0003
Reading a trace from the top tells you what the user saw. Reading down the tree tells you how that result was made, and the duration beside each span tells you where the time went. A slow answer stops being a mystery the first time you see that retrieve_docs took four of the five seconds.
- Pick a small program you have written that has three or four functions calling each other.
- On paper, draw it as a tree: the entry-point function at the top, the functions it calls beneath.
- Decide which of the functions would be a trace and which would be spans.
Threads, datasets, experiments and prompts
Four more nouns complete the picture. You will use all of them by the end of this guide, so a first pass now makes the later sections easier.
A thread groups traces that belong to one conversation. Each message a user sends to a chatbot is its own trace, and you tie them together by giving them a shared thread_id string that you choose. The identifier must be unique within the project. Opik then shows the conversation as a whole, and scoring rules can judge the conversation rather than a single reply. A thread is considered inactive after a cooldown (fifteen minutes by default), after which thread-level scoring can run.
A dataset is a collection of examples you test against. Each item has an input, usually an expected_output, and any other fields you want, such as a category. Datasets in Opik are versioned: every change produces a new immutable version, v1, v2 and so on, and a latest tag points at the newest. Inserting the same item twice does not duplicate it, because Opik deduplicates on insert.
An experiment is the record of one evaluation run. Opik takes each item in a dataset, hands it to your task (the function that runs your application), collects the output, scores it with one or more metrics, and stores everything together: every item, every output, every score and a trace for every item. You can open two experiments side by side to see whether a change helped. One rule is easy to forget: an experiment and all its traces land in the project of the dataset it ran against.
A metric is a scoring function. Heuristic metrics are deterministic and cheap: Equals checks for an exact match, Contains looks for a substring, RegexMatch applies a pattern, IsJson checks validity. LLM-as-a-judge metrics, such as Hallucination and AnswerRelevance, ask another model to grade the output and explain why. Every metric returns a score result with a name, a value and usually a reason.
A prompt in Opik's Prompt Library is a named, versioned template with {{variable}} placeholders. Changing the text creates a new version, and old versions stay available, so you can roll back and see which wording produced which results.
Trace
One request through your app, stored as a tree of spans.
Thread
Several traces that share a conversation id.
Dataset
Versioned examples to test against.
Experiment
One scored run of your task over a dataset.
There is also a newer object in 2.x, the test suite. It is a dataset-like collection where each item carries plain-English assertions ("the answer mentions the 30-day window") that an LLM judge checks. Test suites are covered at the Mid level; for now it is enough to know that they exist and that they suit behaviour you cannot express as an exact match.
- Write three questions and the answers you would accept for a tiny assistant, for example a bot that answers questions about your own university.
- Label each as input or expected_output.
Choosing where Opik runs: Cloud or your laptop
Before installing anything, decide where the server lives. The SDK is only a client. It collects what your code does and sends it over HTTP to a server, and the server stores it and draws the web pages. So there are two separate things to set up: the SDK in your Python environment, and a server it can reach.
Opik Cloud
- Nothing to run or upgrade
- Free tier to start
- Needs an API key and a workspace name
- Data is stored on Comet's infrastructure
- Best for learning and small teams
Local (Docker Compose)
- Runs on your machine, data stays local
- No account, no API key
- Needs Docker and several gigabytes of memory
- Opens at
http://localhost:5173 - Meant for learning, not production
For a first encounter, Cloud is the quickest: sign up at comet.com, create a workspace if prompted, and copy your API key from the user menu at the top right. The local route is better when you are offline, when you want to see what the moving parts are, or when your data must not leave your machine. Both routes are shown below and the rest of the guide works identically on either.
Two things in the local setup surprise newcomers. First, the open-source self-hosted version has no user accounts and no API keys: anyone who can reach the address can use it, so never expose a local instance to the internet. Second, a full local install runs several services (the application backend, a Python backend, the web front end, MySQL, ClickHouse, Redis, ZooKeeper and MinIO object storage), and ClickHouse in particular is hungry. Give Docker Desktop a few gigabytes of memory, or the stack will restart itself in the background.
If you have worked through the Docker guide, the local setup is a good exercise: you will see a real multi-container application start from a single script.
- Decide which route you will use for this guide. If you are unsure, pick Cloud.
- If Cloud: create the account now and copy your API key to a safe place.
- If local: check that
docker --versionandgit --versionboth work.
Installing the SDK and checking it works
The Python SDK is the most complete Opik client and the one this guide uses. It needs Python 3.10 or newer. Always install into a virtual environment so that Opik and its dependencies do not collide with other projects.
python3 -m venv .venv
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install -U opik
opik --version
The last command prints the installed version. You should see something in the 2.2 series. If you see 1.x, you are on an old release and most of this guide will not match, so upgrade with pip install -U opik.
Installing the package gives you two things: the importable library (import opik) and a command-line tool, also called opik, which includes the configure and healthcheck commands you will use in a moment.
Pointing the SDK at a server
The SDK needs to know where the server is and, for Cloud, who you are. The friendliest way is the interactive configuration command.
opik configure
For Cloud it asks for your API key, your workspace and a project name, then writes the answers to a file called ~/.opik.config in your home directory (on Windows, %USERPROFILE%\.opik.config). The project name you give becomes the default for everything you log. Because projects scope your datasets and prompts in Opik 2.x, choosing a meaningful name here, such as support-bot, rather than accepting the default is a good habit.
For a local server, use the --use_local flag, which writes the local address instead of asking for a key.
opik configure --use_local
The resulting file is plain text and looks like this (the exact keys present depend on your answers):
[opik]
url_override = http://localhost:5173/api
workspace = default
project_name = support-bot
For Cloud, the file instead holds api_key and workspace, and the address is the default Cloud URL, https://www.comet.com/opik/api, so no url_override is needed. The four keys you will meet most often are url_override, api_key, workspace and project_name.
You can do the same thing from Python, which is handy in notebooks:
import opik
opik.configure(use_local=False) # Cloud: prompts for your key
# opik.configure(use_local=True) # local server
Environment variables as an alternative
Every key has an environment-variable equivalent. Environment variables win over the file, which makes them the right tool for scripts, containers and continuous integration, where you should not commit a config file.
export OPIK_API_KEY="your-key-here"
export OPIK_WORKSPACE="your-workspace"
export OPIK_PROJECT_NAME="support-bot"
# for a local server instead:
export OPIK_URL_OVERRIDE="http://localhost:5173/api"
On Windows PowerShell the same thing is written $env:OPIK_API_KEY="your-key-here". The precedence, from strongest to weakest, is: arguments you pass in code, then environment variables, then ~/.opik.config, then built-in defaults.
~/.opik.config lives in your home directory, so it is safe by default. A .env file inside your project is not. Add .env to .gitignore before you create it, and never paste an API key into a notebook you will share.
Checking the connection
opik healthcheck
This command inspects your configuration and tries to reach the backend, and it is the first thing to run whenever anything misbehaves. A healthy result reports the URL it is using and that the connection succeeded. If it reports a problem, read the URL first. Nine out of ten beginner failures are a wrong address or a missing key, and the error section later in this guide lists the exact messages.
- Create a virtual environment and install
opik. - Run
opik configure(oropik configure --use_localafter starting the local server in the next section). - Run
opik healthcheckand read the output line by line. - Open
~/.opik.configin an editor and find each key you just set.
Running Opik locally with Docker
Skip this section if you chose Cloud. If you want your own server, the project ships a launcher script that starts the whole stack with Docker Compose. You need git, and Docker with Compose (Docker Desktop provides both on macOS and Windows).
git clone https://github.com/comet-ml/opik.git
cd opik
./opik.sh
On Windows, run the PowerShell version instead:
git clone https://github.com/comet-ml/opik.git
cd opik
powershell -ExecutionPolicy ByPass -c ".\opik.ps1"
The first run downloads several images, so expect a few minutes. When it finishes, the interface is at http://localhost:5173. Open it in a browser and you will see the Opik home page with no projects yet.
The script accepts flags that are worth knowing now, because you will use two of them constantly.
| Flag | What it does |
|---|---|
--verify |
Checks that the containers are healthy |
--stop |
Stops the stack and keeps your data |
--clean |
Stops the stack and deletes all data volumes. Irreversible |
--infra |
Starts only the databases and storage |
--backend |
Starts infrastructure plus the backend |
--guardrails |
Adds the optional guardrails backend |
--build |
Builds images from source |
--help |
Lists every option |
Your data lives in a folder called opik inside your home directory, so stopping and restarting the stack keeps your traces. To upgrade, run ./opik.sh again. To pin a specific release, set the version first: OPIK_VERSION=2.2.86 ./opik.sh.
docker compose down --volumes. Every trace, dataset and experiment you have recorded disappears. Use --stop when you only want to free up memory.
Once the stack is up, connect the SDK to it with opik configure --use_local, then check with opik healthcheck. One point trips up people who use TypeScript: the TypeScript SDK's default address is already the local server (http://localhost:5173/api), while the Python SDK's default is Cloud. The same code therefore behaves differently in the two languages until you configure it explicitly.
This local stack is meant for learning and development. It is not a production deployment. For production, Opik provides a Helm chart for Kubernetes, which is the subject of the Senior level, and connects naturally to the Helm guide and the Kubernetes guide.
- Run
./opik.shand wait for it to finish. - Run
./opik.sh --verifyand read the health report. - Open
http://localhost:5173in your browser. - Run
docker psand count the containers.
Your first traced function
Everything in Opik starts with one decorator. A decorator is a Python feature that wraps a function so that extra behaviour runs around it; you write @name on the line above the function definition. Opik's decorator is track, and it turns each call of the wrapped function into a recorded trace (or a span, if it is called from inside another tracked function).
Create a file called first_trace.py. This example calls no model, so it needs no provider key. That is deliberate: it lets you confirm Opik works before any other variable enters the picture.
import opik
from opik import track
@track
def shout(text: str) -> str:
return text.upper() + "!"
result = shout("hello opik")
print(result)
opik.Opik().flush()
Run it with python first_trace.py. You should see HELLO OPIK! printed, and a few seconds later a new project will appear in the web interface, named after the project you configured (or Default Project). Open it, go to the Logs page, and you will find one trace named shout. Click it. The input panel shows {"text": "hello opik"}, and the output panel shows the returned string. You never told Opik what the input and output were: @track read the function's arguments and return value for you.
The last line deserves an explanation because it is the most common reason beginners see nothing. Opik does not send each record the instant it happens. To keep your application fast, the SDK batches records in memory and sends them from a background thread. In a long-running server that is invisible. In a short script, the program can reach its last line and exit before the background thread has sent anything, and the trace is lost. Calling flush() tells the SDK to send everything that is queued and wait until it is done. You can also flush inside the decorator with @track(flush=True), which is convenient for small scripts and notebooks.
If you want to check quickly that the connection works without writing a file, the documented smoke test is this five-liner, which you can paste in a Python shell:
from opik import track
import opik
@track
def f(x):
return x
f("hi")
opik.Opik().flush()
- Save the file above and run it.
- Find the trace in the Logs page of your project.
- Delete the
flush()line, run the script again and see whether the new trace appears. - Put the line back, then try
@track(flush=True)instead.
Nested spans: seeing the tree
A single traced function is a trace with no children. The power appears when tracked functions call each other. The rule is simple: the outermost tracked call becomes the trace, and every tracked call made inside it becomes a span nested under its caller. You never wire the parent-child links yourself.
import opik
from opik import track
@track(name="load-policy")
def load_policy(topic: str) -> str:
policies = {"refund": "Refunds are accepted within 30 days of purchase."}
return policies.get(topic, "No policy found.")
@track(name="write-answer")
def write_answer(question: str, policy: str) -> str:
return f"Q: {question} | A: {policy}"
@track(name="answer-question")
def answer_question(question: str) -> str:
policy = load_policy("refund")
return write_answer(question, policy)
print(answer_question("How long do I have to return an item?"))
opik.Opik().flush()
In the interface, the Logs page shows one trace named answer-question. Open it and you see the tree in the left panel: answer-question at the root, load-policy and write-answer as its children. Clicking a span shows that span's own input and output, and each has a duration. This is the same structure as the conceptual diagram earlier, now produced by two lines of decoration.
Notice the name= argument. By default the span name is the function name, which is fine, but explicit names let you rename a function later without breaking the way you search and filter old traces.
The decorator accepts more options. The ones you will use in this first month are:
| Option | Purpose |
|---|---|
name="..." |
The display name of the trace or span |
type="llm" |
Marks the span as a model call, which enables cost tracking (other values: general, tool, guardrail) |
project_name="..." |
Sends this trace to a specific project |
flush=True |
Sends immediately after the function returns |
environment="production" |
Tags the trace with a lifecycle label |
If a tracked function raises an exception, Opik records the failure on the span and then lets the exception continue as normal. A failing step therefore shows up in the tree exactly where it failed, which is one of the most useful things tracing gives you.
- Run
nested.pyand open the trace. - Add a third tracked function that is called from
load_policy, then run again. - Make
write_answerraiseValueError("boom")and run once more.
Tracing a real model call
Now add an actual model. You have two routes: wrap the provider's client so every call is traced automatically, or mark your own function as a model call by hand. Start with the wrapper, because it gives you model name, token counts and cost with no extra effort.
Opik has integrations for dozens of providers and frameworks, including OpenAI, Anthropic, Amazon Bedrock, Gemini, Ollama, LiteLLM, LangChain, LangGraph, LlamaIndex, CrewAI and the OpenAI Agents SDK. The pattern is nearly always the same: import a helper from opik.integrations, wrap your client or attach a callback, and carry on. Here it is for OpenAI.
pip install openai
export OPENAI_API_KEY="sk-..."
import opik
from openai import OpenAI
from opik import track
from opik.integrations.openai import track_openai
client = track_openai(OpenAI())
@track(name="ask-model")
def ask(question: str) -> str:
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": "Answer in one short sentence."},
{"role": "user", "content": question},
],
)
return response.choices[0].message.content
print(ask("What does an LLM trace record?"))
opik.Opik().flush()
The wrapped client behaves exactly like the normal one. The only difference is that each chat.completions.create call now emits a span of type llm, nested under ask-model because ask is tracked. Open the trace and click the inner span. You will see the full list of messages sent, the model's reply, the model name, prompt and completion token counts and an estimated cost in dollars. The cost comes from a built-in price table keyed on the model, so a very new or custom model may show no cost rather than a wrong one.
Two practical points. First, track_openai takes an optional provider= argument if you point the OpenAI client at a compatible service, such as a gateway or a local server, so the span records the right provider. Second, you do not have to use the decorator at all: a wrapped client on its own produces a trace per call, but the decorator is what lets you group several calls into one meaningful trace.
If you use LangChain or LangGraph, the integration is a callback object. For LangGraph the current pattern is to wrap the compiled graph:
from opik.integrations.langchain import OpikTracer, track_langgraph
app = track_langgraph(app, OpikTracer())
An older pattern, where you pass OpikTracer(graph=app.get_graph(xray=True)) and add config={"callbacks": [tracer]} to each call, is still common in tutorials and still works, but the wrapper form is simpler. If you are learning LangGraph itself, see the LangGraph guide.
Doing it by hand when no integration fits
Sometimes you call a model through a library Opik has no wrapper for. You can still get full cost tracking by telling Opik what the span is. Put type="llm" on the decorator and update the span with the model, provider and usage.
import opik
from opik import track, opik_context
@track(name="custom-model-call", type="llm")
def call_my_model(prompt: str) -> str:
answer = "stub answer" # replace with a real call
opik_context.update_current_span(
model="my-model",
provider="custom",
usage={"prompt_tokens": 20, "completion_tokens": 5, "total_tokens": 25},
)
return answer
call_my_model("Hello")
opik.Opik().flush()
Cost is only calculated when the span is an llm span and carries a model, a provider and usage. Miss one of the three and the cost column stays empty. If the model is not in Opik's price table, cost is reported as missing, which is honest and safer than a guess.
- Run
llm_call.pywith your own key and find the trace. - Click the inner LLM span and read the model, token counts and cost.
- Change the question and the model name, run again, and compare the two traces.
manual_llm.py instead and confirm the usage numbers you supplied appear on the span.
Adding context: tags, metadata, threads and feedback
A trace that only contains input and output is useful. A trace you can search, group and score is far more useful. Opik gives you four ways to attach extra information, and all of them are set from inside a tracked function through opik_context.
Tags are short labels, used for filtering: "support", "beta-user", "experiment-b". Metadata is a dictionary of arbitrary details: the model version, a customer tier, the git commit, the name of the retriever. Use tags for coarse categories you will filter by constantly, and metadata for everything else. A thread id groups traces into a conversation. Feedback scores are named numbers, optionally with a reason, recording a judgement: did the user like it, was it correct, was the tone right.
import opik
from opik import track, opik_context
@track(name="chat-turn")
def chat_turn(user_id: str, conversation_id: str, message: str) -> str:
reply = f"You said: {message}"
opik_context.update_current_trace(
tags=["support", "demo"],
metadata={"user_tier": "free", "user_id": user_id},
thread_id=conversation_id,
feedback_scores=[{"name": "user_feedback", "value": 1, "reason": "thumbs up"}],
)
return reply
chat_turn("u-42", "conv-001", "Hi there")
chat_turn("u-42", "conv-001", "What are your hours?")
opik.Opik().flush()
After this runs, open the project. The two calls are two separate traces, but they share the thread id conv-001, so the Threads view shows them as one two-message conversation. The tags appear as coloured chips you can filter on, the metadata is listed in the trace details, and user_feedback shows as a score column. You can filter the Logs page by any of those, for example to see only traces tagged support whose user_feedback is below one.
There is an important subtlety in the function names. update_current_trace changes the trace at the top of the current tree, while update_current_span changes only the span you are currently inside. When you want to record which retriever a particular step used, that belongs on the span. When you want to record which customer the whole request belonged to, that belongs on the trace.
opik_context.update_current_span(metadata={"retriever": "bm25", "top_k": 5})
The decorator can also receive this context at call time, without changing the function body, through an opik_args dictionary. That is useful when a framework calls your function for you and you cannot edit the inside.
Scores can also be added in the interface. Open a trace, find the feedback-scores area and add a score by hand, which is how human reviewers mark answers "good" or "bad". Names for human scores can be declared ahead of time as feedback definitions, so everyone uses the same vocabulary and the same scale. Later, the same score names will be written automatically by your evaluation metrics and by online scoring rules, so a single column can mix human and automatic judgement.
thread_id on every turn, and it must be unique per conversation within the project. Use your own session or conversation identifier rather than a random value generated on each call, or every message becomes a one-message thread.
- Run
context.py. - In the project, open the Threads view and find
conv-001. - Filter the Logs page by the tag
support. - Add a manual feedback score to one trace in the interface.
Projects and environments
Since Opik 2.x makes the project the home of everything an application owns, it pays to understand exactly how the SDK decides which project a trace goes to. The answer is an ordered list, and the first match wins:
- An explicit argument, such as
@track(project_name="support-bot")oropik.Opik(project_name="support-bot"). - The environment variable
OPIK_PROJECT_NAME. - The
project_nameline in~/.opik.config. - The literal string
Default Project.
If nothing is configured, your traces land in Default Project and the SDK logs a one-time warning. That is the usual reason people "lose" traces: they are there, just in a different project than the one they are viewing. Decide your project at the top of the program and keep it there.
import opik
from opik import track
client = opik.Opik(project_name="support-bot")
@track(project_name="support-bot")
def handle(question: str) -> str:
return question[::-1]
handle("hello")
client.flush()
There is one rule that surprises people. If you wrap code in opik.project_context("name"), the outermost project context wins. A nested @track(project_name="other") inside it is ignored, with a warning, by design, so that one trace never gets split across projects. Set the project at the outermost level and do not rely on inner overrides.
Environments are a second, lighter label. Where a project separates applications, an environment separates the stages of one application: development, staging, production. You can set it per call with @track(environment="production"), or globally with the OPIK_ENVIRONMENT variable, and filter the Logs page by it. Keeping staging and production in the same project, separated by an environment label, makes comparisons easy. Keeping them in separate projects gives cleaner isolation. Either is reasonable; what matters is to choose one and be consistent.
- Run
projects.pywithOPIK_PROJECT_NAMEset to a different name. - Check which project the trace landed in, then remove the
project_nameargument and run again.
Datasets: your test questions, saved
Tracing tells you what happened. Evaluation asks whether it was good. That needs a fixed set of examples so that you can run the same questions before and after every change, and in Opik that set is a dataset.
You create and fetch datasets through the client. get_or_create_dataset is the safe form: it returns the dataset if it exists and creates it if not, so the script can be run repeatedly.
import opik
client = opik.Opik(project_name="support-bot")
dataset = client.get_or_create_dataset(name="refund-questions", project_name="support-bot")
dataset.insert([
{"input": "How long do I have to return an item?", "expected_output": "30 days"},
{"input": "Can I get a refund after 45 days?", "expected_output": "No"},
{"input": "Do refunds go to the original payment method?", "expected_output": "Yes"},
])
print("items:", len(dataset.get_items()))
Each item is a plain dictionary. Only the field names you choose matter, and later your task and metrics refer to them by name. Keep input and expected_output as conventions, since many built-in metrics look for those names, but nothing stops you adding category, difficulty or a source_document.
Run the script twice. The second run does not duplicate the items, because Opik deduplicates on insert. Each change produces a new immutable version of the dataset, and the interface lets you browse versions, so you can always tell which data an old experiment used. In the interface you can also add, edit and delete items in a draft, and committing the draft creates a new version.
Datasets can be loaded from files too. The client offers helpers such as insert_from_pandas and read_jsonl_from_file, which matter once your examples live in a spreadsheet or a JSONL export rather than in code.
- Run
make_dataset.py, then open the project's Datasets page. - Open
refund-questionsand look at the items and the version list. - Run the script a second time and confirm the item count does not change.
Running your first experiment
An experiment ties the pieces together: a dataset, a task that produces an answer for each item, and metrics that score each answer. The function that does this is evaluate.
The task is an ordinary Python function. Opik calls it once per dataset item, passing the item as a dictionary, and expects a dictionary back. The key names in the returned dictionary are what your metrics will read.
The example below uses a stand-in model so you can run it without a provider key. Swap in a real call later.
import opik
from opik.evaluation import evaluate
from opik.evaluation.metrics import Contains
client = opik.Opik(project_name="support-bot")
dataset = client.get_or_create_dataset(name="refund-questions", project_name="support-bot")
def task(item: dict) -> dict:
question = item["input"]
# Stand-in for a model call. Replace with your real application.
if "return an item" in question:
answer = "You have 30 days to return an item."
elif "45 days" in question:
answer = "No, that is outside the window."
else:
answer = "Yes, to the original payment method."
return {"output": answer}
result = evaluate(
dataset=dataset,
task=task,
scoring_metrics=[Contains(name="contains_expected")],
scoring_key_mapping={"reference": "expected_output"},
experiment_name="baseline-v1",
project_name="support-bot",
experiment_config={"model": "stand-in", "prompt_version": "none"},
)
When it finishes, Opik prints a summary and a link to the experiment. Under the hood, it ran your task on every item (using a pool of worker threads, eight by default), sent each output to each metric, and stored the results. Open the experiment page and you will see a table with one row per dataset item: the input, the expected output, what your task returned, and a score column for each metric.
The Contains metric checks whether a reference text appears inside the output. Its score method takes the task's output and a reference. Opik passes each metric the fields of the dataset item plus the task's returned dictionary, matched by name. Your dataset calls the field expected_output, but this metric asks for reference, so the scoring_key_mapping argument bridges the two: {"reference": "expected_output"} means "fill the metric's reference argument from the item's expected_output field". The same argument works for any renaming, for instance {"input": "user_question"} when your dataset calls the question field user_question. Mapping is the single most common thing beginners must learn about metrics, and a missing one shows up as an error about a missing argument.
The third question expects Yes and the stand-in returns Yes, ..., so it scores 1. The second expects No and the answer begins No,, so it also scores 1. Whether you care about case-sensitivity, exact match or a looser judgement is a choice you make by picking the metric, and this is the main design decision in evaluation.
Three other arguments from this call matter. experiment_name is the label you will search by later, so make it descriptive, for example baseline-v1 and then shorter-prompt-v2. experiment_config is a dictionary of whatever describes this run (model, prompt version, temperature), stored with the experiment so that a month from now you can tell what produced the numbers. project_name should match your dataset's project. Remember the rule from earlier: an experiment lands in the project of its dataset.
One limit to know early: evaluate does not accept an async task function. Write the task as an ordinary synchronous function; Opik provides the concurrency for you.
- Run
run_experiment.pyand open the experiment from the printed link. - Change the stand-in answer for one question so it is wrong, then run again with
experiment_name="baseline-v2". - Open the Experiments page, select both runs and use the compare view.
Choosing metrics: heuristics and LLM judges
The metric decides what "good" means, so it deserves more thought than any other argument. Opik ships a large library, and they fall into three families.
Heuristic metrics are deterministic functions of the output. They are fast, free and perfectly repeatable, and they should be your first choice whenever the thing you care about can be checked by a rule. Equals demands an exact match. Contains looks for a substring. RegexMatch tests a pattern, useful for phone numbers or identifiers. IsJson checks that the output parses as JSON, which is invaluable for applications that must emit structured data. String-similarity metrics such as Levenshtein, BLEU, ROUGE, ChrF and BERTScore compare the output with a reference text when an exact match is too strict.
LLM-as-a-judge metrics ask another model to grade the output. They handle qualities no rule can capture: Hallucination asks whether the answer contains claims not supported by the provided context, AnswerRelevance asks whether it addresses the question, ContextPrecision and ContextRecall judge retrieval quality, Moderation flags unsafe content, and GEval lets you describe a custom criterion in plain English. Each returns a score between 0 and 1 and usually a written reason, which is where much of the debugging value lies.
Conversation metrics score a whole thread rather than one reply, for example whether a dialogue stayed coherent. They are covered at the Mid level.
from opik.evaluation.metrics import Hallucination
metric = Hallucination()
result = metric.score(
input="How long is the refund window?",
output="The refund window is 90 days.",
context=["Refunds are accepted within 30 days of purchase."],
)
print(result.value, result.reason)
This needs a model to do the judging. By default the Python SDK uses openai/gpt-5-nano, configurable through the OPIK_DEFAULT_LLM environment variable, so an OPENAI_API_KEY must be set. To use a different judge, pass a model string in LiteLLM format, such as Hallucination(model="bedrock/..."). The judge reads your text, so anything in it leaves your machine for the provider. For regulated data, choose a provider and region that your employer permits, or a locally hosted model.
The output of a score is a ScoreResult with a name, a value and a reason. For the hallucination metric, a higher value means more hallucination, so read each metric's documentation before you interpret the direction. This is a frequent source of confused charts.
Use a heuristic when
- The answer has a known right form
- You need it free and repeatable
- You run it on every commit
- Checking format, presence or pattern
Use a judge when
- Quality is a matter of meaning
- You can afford provider tokens
- You want a reason with each score
- Checking faithfulness or relevance
A sensible starting combination is one or two cheap heuristics plus one judge. Use the heuristics to catch the broken cases cheaply, and the judge to watch the subtle quality. Judges are themselves imperfect and sometimes wrong, so read a sample of their reasons before you trust the averages. Adding a judge to the earlier experiment is a one-line change:
from opik.evaluation.metrics import Contains, Hallucination
scoring_metrics = [Contains(name="contains_expected"), Hallucination()]
Because the judge also needs the retrieved context, your task must return it under the key the metric expects, for example {"output": answer, "context": [doc1, doc2]}. When a metric complains about a missing argument, it is telling you exactly which key the task should return or the mapping should supply.
BaseMetric and implement score, returning a ScoreResult. Include **ignored_kwargs in the signature so the metric tolerates extra fields. Custom metrics are covered in depth at the Mid level.
- Add
RegexMatchorIsJsonto your experiment and decide what pattern would make sense for one of your questions. - If you have a provider key, run the
Hallucinationsnippet above and read thereason. - Change the
outputso it agrees with the context and run it again.
Managing prompts in the Prompt Library
A prompt is code that happens to be written in English. It changes often, it affects behaviour, and a silent change can make quality drop. So it should be versioned like code. Opik's Prompt Library stores named, versioned prompt templates with {{variable}} placeholders, and lets your application fetch a specific version at runtime.
import opik
client = opik.Opik(project_name="support-bot")
prompt = client.create_prompt(
name="refund-answerer",
prompt="You are a support agent. Answer in one sentence.\nPolicy: {{policy}}\nQuestion: {{question}}",
project_name="support-bot",
)
print(prompt.format(
policy="Refunds are accepted within 30 days.",
question="Can I return this?",
))
create_prompt is careful about versions. If a prompt with that name exists and the text is identical, nothing new is created. If the text changed, you get a new version, v2, v3 and so on; old versions are immutable. prompt.format(...) fills the placeholders and returns the final string, ready for a model call. To fetch an existing prompt, including a particular version:
latest = client.get_prompt(name="refund-answerer", project_name="support-bot")
pinned = client.get_prompt(name="refund-answerer", version="v1", project_name="support-bot")
There is a bonus that makes this worth doing from the start. When you call get_prompt inside a tracked function, Opik links that exact prompt version to the trace. Later, when one trace looks bad, you can see which wording produced it, and compare traces across versions. You can also pass prompts=[prompt] into evaluate so that an experiment records which prompt version it tested.
Prompts come in two kinds. A text prompt is a single string, as above. A chat prompt is a list of role-tagged messages (system, user, assistant) and uses create_chat_prompt and get_chat_prompt. Prompts are fetched and cached by the SDK, so edits made in the interface reach running code after a short delay (five minutes by default, controlled by OPIK_PROMPT_CACHE_TTL_SECONDS). That is helpful to know if an edit seems not to take effect.
Finally, prompts can be edited in the interface, and Opik's Playground lets you try a prompt against a model and even run it over a dataset, using provider keys you configure once per workspace under Configuration and AI Providers.
- Run
prompts.pyand find the prompt in the Prompt Library in the interface. - Change one word in the template and run the script again.
- Open the prompt and look at the version list.
Reading the interface
The code you have written produces records, but you spend much of your time reading them, so it is worth learning the layout of the web interface. The names below match Opik 2.x.
The left navigation lists the areas of a project. Logs is the page you will use most: threads, traces and spans are all shown on this one page, with tabs to switch between them. Above the table are a search box, column controls and filters. Clicking a row opens a panel with the span tree on the left and the details on the right: input, output, metadata, feedback scores and usage. Experiments no longer have their own trace list elsewhere: the traces from an experiment appear as a Logs tab on the experiment page.
Dashboards and metrics show cost, token usage, latency and error counts over time, filtered by project. This is where the price of your application becomes visible. Datasets lists your collections and their versions. Experiments lists the runs, with a compare mode. Prompts is the library. Annotation queues let you assemble a batch of traces for human reviewers to score. Optimization runs (formerly "Optimization studio") manage automated prompt improvement. Rules hold online scoring, where a judge metric automatically scores a sample of your live traces.
Filtering uses a small language. In the filter bar, conditions look like <COLUMN> <OPERATOR> <VALUE>. Strings need double quotes, nested values use dots such as metadata.model or feedback_scores.user_feedback, and conditions are joined with AND. The language has no OR, which surprises people: to see either of two cases, run two filters.
There is also a built-in assistant called Ollie, which can read your traces and help you find patterns. It is a convenience layer, and it uses credits on Cloud. You do not need it to follow this guide.
One last interface trick is useful to know. Press Cmd+Shift+. on macOS, or Ctrl+Shift+. on Windows and Linux, to toggle Comet Debugger Mode, which displays the server version and the round-trip time to it. When a bug report is needed, or when you suspect the SDK and the server are at mismatched versions, this is the quickest way to read the server's version.
- On the Logs page, switch between the Threads, Traces and Spans tabs.
- Add a filter such as
tags contains "support"or one on a feedback score. - Find the cost and token charts on the metrics page.
A short look at TypeScript and other languages
Everything so far used Python. If you write JavaScript or TypeScript, Opik has an SDK for Node, installed with npm install opik, and a setup command, npx opik-ts configure, which needs Node 18.17 or newer. Its concepts are identical, and its calls mirror the Python low-level client.
import { Opik } from "opik";
const client = new Opik({ projectName: "support-bot" });
const trace = client.trace({
name: "answer-question",
input: { question: "How long is the refund window?" },
output: { answer: "30 days" },
});
trace.span({
name: "call-model",
type: "llm",
input: { prompt: "..." },
output: { text: "30 days" },
});
trace.end();
await client.flush();
There is no decorator here; instead you create the trace and its spans explicitly, end the trace and flush. Integrations ship as separate npm packages, for example opik-openai whose trackOpenAI wraps the OpenAI client and opik-vercel for the Vercel AI SDK.
For any other language, Opik offers a REST API and OpenTelemetry support: a standard way of exporting trace data that many languages already speak. Point an OpenTelemetry exporter at the Opik endpoint, https://www.comet.com/opik/api/v1/private/otel on Cloud, or http://localhost:5173/api/v1/private/otel locally, over HTTP (gRPC is not supported). Headers carry your key, project and workspace. The Mid level covers this properly.
- If you use Node, install
opik, runnpx opik-ts configureand run the snippet above. - Otherwise, read the TypeScript SDK page on the official site and note where its names differ from Python.
Configuration you will actually touch
You have now met most of the configuration surface. This section gathers it in one place so you can look it up.
| Setting | Config file key | Environment variable | Notes |
|---|---|---|---|
| Server address | url_override |
OPIK_URL_OVERRIDE |
Cloud default; local is http://localhost:5173/api |
| API key | api_key |
OPIK_API_KEY |
Cloud only; self-hosted open source has none |
| Workspace | workspace |
OPIK_WORKSPACE |
default on a self-hosted install |
| Project | project_name |
OPIK_PROJECT_NAME |
Falls back to Default Project |
| Environment | (none) | OPIK_ENVIRONMENT |
Lifecycle label such as production |
| Switch tracing off | opik_track_disable |
OPIK_TRACK_DISABLE |
Also opik.set_tracing_active(False) |
| Judge model | (none) | OPIK_DEFAULT_LLM |
Default openai/gpt-5-nano |
| Usage analytics | analytics_enable |
OPIK_ANALYTICS_ENABLE |
On by default in the Python SDK; set to false to opt out |
Some observations on that table. The config file is the friendliest place for values that belong to you personally and never change, and environment variables are better for anything a script or container needs to override. The OPIK_TRACK_DISABLE switch is worth knowing: in unit tests you usually do not want your code sending traces, so set it to true in the test environment, or call opik.set_tracing_active(False).
Older material sometimes uses OPIK_BASE_URL for the server address. The current configuration reference uses OPIK_URL_OVERRIDE, so use that one.
Logging deserves one warning. To see what the SDK itself is doing while debugging, use the environment variables OPIK_CONSOLE_LOGGING_LEVEL (default INFO) and OPIK_FILE_LOGGING_LEVEL, and set them before your program imports opik. The usual Python approach, logging.getLogger("opik").setLevel(...), has no effect, because the SDK manages its own logging. Likewise, if you load secrets from a .env file with python-dotenv, call load_dotenv() before import opik, or the values will be read too late and ignored.
from dotenv import load_dotenv
load_dotenv() # first
import opik # then import opik
- Run a traced script with
OPIK_CONSOLE_LOGGING_LEVEL=DEBUGset and look at the extra output. - Set
OPIK_TRACK_DISABLE=trueand run it again.
When things go wrong: common errors
Most of the trouble you will have falls into a handful of patterns. The format here is the message you will see or the symptom, what causes it, and the fix.
Nothing appears in the interface. The most frequent cause in short scripts is that the program exited before the background sender delivered the batch. Call opik.Opik().flush() at the end, or use @track(flush=True). If flushing does not help, run opik healthcheck and compare the URL, workspace and project it prints with the ones in your browser.
Traces appear in Default Project. You did not configure a project, or a project context overrode your intention. Set OPIK_PROJECT_NAME, pass project_name, or rerun opik configure. A one-time warning in the console says this happened.
HTTP 403 Forbidden. The API key is wrong or missing, you are pointing at a workspace you cannot access, or the URL is wrong. Re-run opik configure and check OPIK_API_KEY, OPIK_WORKSPACE and OPIK_URL_OVERRIDE. For a self-hosted server the URL must end in :5173/api; the /api is because the front end's web server forwards that path to the backend.
A certificate error.
[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: self-signed certificate in certificate chain (_ssl.c:1006)
This appears when your company network inspects HTTPS traffic with its own certificate, or when a server uses a self-signed one. The correct fix is to tell Python to trust your organisation's certificate authority by setting REQUESTS_CA_BUNDLE to the path of its certificate bundle. Setting OPIK_CHECK_TLS_CERTIFICATE=false also makes the error disappear, but it turns off verification, so treat it as a short-lived last resort rather than a solution.
HTTP 429 and a log line such as OPIK: Ingestion rate limited, retrying in 55 seconds. Opik Cloud limits how fast you may send. The SDK waits and retries automatically, so the message is informational. If you hit it constantly, reduce how many threads you run in parallel.
HTTP 413, payload too large. A single field was enormous, typically a full document or a base64 image inside an input or output. Log identifiers, summaries or the top few results instead of entire blobs.
Updates lost and a batching warning. If you use the lower-level client and call .end() or .update() on a trace immediately after creating it, the update can arrive before the create and be dropped. The decorator and context managers avoid the problem. This matters more at the Mid level.
A metric fails with an unexpected-argument or missing-argument error. The keys your task returns and the dataset fields do not match what the metric's score method expects. Print what the task returns, read the metric's documentation for its argument names, and add a scoring_key_mapping.
An evaluation refuses an async task. evaluate does not support async def tasks. Make the task a normal function.
Datasets, prompts or experiments seem to vanish after upgrading from 1.x on a self-hosted server. The 2.0 release made these project-scoped, and old items without a project are hidden until migration jobs run. On Cloud the migration happened automatically. On self-hosted the procedure is in the official migration guide and is a Senior-level topic, but the symptom is worth recognising.
uvx: command not found when following MCP instructions. That is the uv tool missing, or a terminal opened before installing it. Install uv, then open a new terminal.
opik healthcheck. Then confirm the URL, workspace and project. Then check the script flushes. Then turn on OPIK_CONSOLE_LOGGING_LEVEL=DEBUG. Only after all four should you suspect a bug in Opik itself.
- Set
OPIK_URL_OVERRIDEto a wrong address such ashttp://localhost:9999/apiand runopik healthcheck. - Read the error, then restore the correct value and run it again.
Putting it all together
Here is one small end-to-end project that uses everything above: a tiny question-answering assistant that is traced, has a versioned prompt, is tested against a dataset, and is scored. It has no external dependency beyond Opik so you can run it anywhere, and a marked spot where you can drop in a real model.
Create a folder called refund-bot and add these files. First, the application.
import opik
from opik import track, opik_context
PROJECT = "refund-bot"
client = opik.Opik(project_name=PROJECT)
POLICY = "Refunds are accepted within 30 days of purchase. Refunds go to the original payment method."
prompt = client.create_prompt(
name="refund-answerer",
prompt="Policy: {{policy}}\nQuestion: {{question}}\nAnswer in one sentence.",
project_name=PROJECT,
)
@track(name="retrieve-policy")
def retrieve_policy(question: str) -> str:
return POLICY
@track(name="generate", type="llm")
def generate(filled_prompt: str) -> str:
# Replace this stub with a real model call, for example via track_openai.
text = filled_prompt.lower()
if "how long" in text or "return" in text:
return "You have 30 days to return an item."
if "45 days" in text:
return "No, that is outside the 30 day window."
return "Yes, refunds go to the original payment method."
@track(name="answer-question", project_name=PROJECT)
def answer_question(question: str, conversation_id: str = "demo") -> str:
policy = retrieve_policy(question)
filled = prompt.format(policy=policy, question=question)
answer = generate(filled)
opik_context.update_current_trace(
tags=["refund-bot"], thread_id=conversation_id, metadata={"prompt_version": "latest"}
)
return answer
Then the evaluation script, which imports the application and scores it against a dataset.
import opik
from opik.evaluation import evaluate
from opik.evaluation.metrics import Contains
from app import PROJECT, answer_question
client = opik.Opik(project_name=PROJECT)
dataset = client.get_or_create_dataset(name="refund-questions", project_name=PROJECT)
dataset.insert([
{"input": "How long do I have to return an item?", "expected_output": "30 days"},
{"input": "Can I get a refund after 45 days?", "expected_output": "No"},
{"input": "Do refunds go to the original payment method?", "expected_output": "original payment method"},
])
def task(item: dict) -> dict:
return {"output": answer_question(item["input"])}
evaluate(
dataset=dataset,
task=task,
scoring_metrics=[Contains(name="contains_expected")],
scoring_key_mapping={"reference": "expected_output"},
experiment_name="refund-bot-baseline",
project_name=PROJECT,
experiment_config={"prompt": "refund-answerer", "model": "stub"},
)
client.flush()
Run it from inside the folder:
cd refund-bot
python evaluate.py
What you should see and what to do with it. The console prints an evaluation summary and a link. In the interface, the refund-bot project now contains one experiment, refund-bot-baseline, with three rows. Open any row's trace from the experiment's Logs tab and you will see the tree you built: answer-question at the top with retrieve-policy and generate beneath it, and the generate span marked as an LLM span. The Prompts page shows refund-answerer. The Datasets page shows refund-questions. Every object belongs to one project, as Opik 2.x intends.
Now make a change and watch it register. Edit the prompt text in app.py, or change the stub's answer for the first question to something wrong, and run python evaluate.py again with a new experiment_name, for example refund-bot-edit-1. Open the Experiments page, select both runs and compare. You have just done what evaluation-driven development means: propose a change, measure it, and decide from data.
Finally, swap the stub for a real model. Use track_openai as shown earlier, call the model inside generate, and mark it with type="llm". Add Hallucination() to scoring_metrics and return "context": [POLICY] from the task. Everything else stays the same, which is the point of the design: the harness does not care what is inside generate.
- Build the
refund-botfolder and runpython evaluate.py. - Find the experiment, its trace tree, the prompt and the dataset in the interface.
- Break one answer on purpose, rerun under a new experiment name and compare the two runs.
- Add a fourth dataset item that the bot should refuse, and see how the bot fares.
What you can now do, and what comes next
If you have worked through the Try-it tasks, you can now do things that many teams skip. You can explain the difference between a workspace, a project, a trace and a span without hesitating. You can install the SDK, point it at Cloud or a local server, and verify the connection with opik healthcheck. You can trace any Python function with @track, nest functions to get a tree, wrap a model client to capture tokens and cost, and attach tags, metadata, thread ids and feedback scores. You can build a versioned dataset, run evaluate with a task and metrics, compare two experiments, and keep your prompts in the Prompt Library. And you know the usual reasons for silence: the missing flush, the wrong project, the wrong URL.
Keep these habits from day one. Set a project name explicitly, and never rely on Default Project. Flush in every script. Name experiments so that you can read the list a month later. Put the model, the prompt version and any other settings into experiment_config. Keep secrets out of function arguments. Add a failing real-world example to your dataset every time production surprises you.
The Mid-level guide starts where this one stops. It explains how batching and the client work precisely enough to predict their behaviour, covers the low-level client and distributed tracing across services, custom metrics and test suites with natural-language assertions, online evaluation of production traffic, annotation queues, the command-line export and import tools, OpenTelemetry, and integration into continuous delivery. The Senior guide treats Opik as a platform for a team: its architecture, the ClickHouse-centred scaling limits, security and access control, upgrade and migration procedures such as the 1.x to 2.x move, backup, cost control and incident playbooks.
Neighbouring tools are worth knowing as you grow. Langfuse solves an overlapping problem with a different philosophy, RAGAS provides retrieval-specific metrics, and LangGraph is a common framework whose runs you will want to trace. The MLflow guide shows how classical machine-learning experiment tracking compares.
Sources
- Opik documentation home
- Quickstart
- Tracing concepts
- Getting started with tracing
- Log traces
- Log chat conversations
- SDK configuration
- Cost tracking
- Annotate traces
- Evaluation concepts
- Evaluation getting started
- Evaluate your LLM application
- Manage datasets
- Metrics overview
- Prompt Library
- Local deployment
- Self-hosting overview
- Opik 2.0 upgrade guide
- TypeScript SDK
- OpenTelemetry integration
- Changelog
- Opik on GitHub