تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
OpikLLMOpsObservability & tracing3 مستويات124 قسمًايغطّي Opik 2.2دليل بالإنجليزية

The Complete Opik Guide

Trace, evaluate and optimise LLM apps with Comet’s open-source Opik. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
23sections
32examples

This is part one of three. It covers everything you need to start doing real work with Opik, not a teaser. By the end you will have installed the SDK, pointed it at either Opik Cloud or a server on your own laptop, traced a small LLM application so that every step shows up in a web page, attached tags and scores to what you traced, built a dataset of test questions, and run an experiment that scores your application against it. Mid-level and Senior take the same topics further; nothing here is thrown away.

This guide is written against Opik 2.2.86, released on 30 September 2026. Opik ships several patch releases a week, and the server, the Python SDK, the TypeScript SDK and the Helm chart all share one version number. So treat "2.2.x" as the version you are learning and run pip install -U opik whenever you start a new project. Where something changed recently and old tutorials will mislead you, this guide says so.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and these ideas only stick once you have watched your own trace appear in the interface and then broken it on purpose.

What Opik is, and the problem it solves

Opik is an open-source platform (Apache-2.0 licence) from Comet for watching, testing and improving applications built on large language models. An LLM application is any program that sends text to a model such as GPT, Claude or Gemini and does something with the answer: a chatbot, a document summariser, a retrieval-augmented question-answering service, an agent that calls tools. Opik records what those programs do, lets you score how well they did it, and keeps the prompts and test data you use to make them better.

YOUR APPcalls an LLM
→
OPIK SDKrecords each step
→
OPIK SERVERstores and scores
→
WEB UIyou read the story

The reason such a tool exists is that LLM applications fail in a way ordinary software does not. A normal function either returns the right number or it does not, and a unit test can tell you which. An LLM application returns fluent, confident text every time, and whether that text is correct, relevant, safe and grounded in your documents is a judgement call. A program can run without a single exception for a month while quietly telling customers something false. Nothing crashes, so nothing alerts you.

Opik attacks that problem from three directions, and these three directions are the structure of the rest of this guide.

Observability means being able to see what happened. Every request through your application becomes a record you can open: the question, the prompt that was actually sent, the documents retrieved, the model's answer, how many tokens it used, what it cost, and how long each step took. When a user says "the bot gave me a weird answer yesterday", you search for it and read the exact chain of events rather than guessing.

Evaluation means being able to measure whether the application is any good. You collect a set of example questions, run your application over all of them, and score the answers automatically. When you change the prompt or swap the model, you run the same set again and compare numbers instead of trusting a feeling.

Improvement and protection cover the rest: a library where prompts are versioned like code, annotation tools so a human can review real answers, online scoring that checks production traffic as it arrives, alerts that message your team when error rates or costs spike, and optimisers that propose better prompts. Some of these (guardrails that block unsafe output inline, for instance) are enterprise features, and this guide only mentions them so that you know they exist.

Opik is available in three forms, and it is worth understanding the difference early because it changes which steps you follow. Opik Cloud is hosted by Comet, has a free tier, and needs nothing installed except the SDK. The self-hosted version is the same open-source code running on your own machine or cluster; the local Docker Compose setup is ideal for learning, and Kubernetes with Helm is the production route. Comet Enterprise adds single sign-on, fine-grained roles and similar controls. The SDK you write code against is identical in all three. Only the address it talks to changes.

For readers in the Gulf and Egypt, the choice between Cloud and self-hosted is often a data-residency question rather than a technical one. If your employer cannot send customer conversations to a vendor's region, a self-hosted Opik inside your own cloud account or data centre keeps prompts and outputs within your boundary. You will see where that choice is made in one configuration line.

Try it
  1. Think of an LLM feature you have used or built: a chatbot, a summary button, a search box with an AI answer.
  2. Write down three ways it could give a wrong answer without raising an error.
  3. For each, write what you would need to see in order to find out why.
three silent failures and a list of things you would want recorded, such as the exact prompt, the retrieved context, the model name and the latency. That list is, almost item for item, what Opik captures for you.

What people did before

Before dedicated tooling, most teams handled LLM applications with print statements, log files and a spreadsheet. The habit is understandable, and it works for the first afternoon. It stops working when the application grows, and seeing where it stops explains why Opik is designed the way it is.

A print of the model's answer tells you the final text but not the intermediate steps. Modern applications are rarely a single call. A retrieval-augmented generation (RAG) service embeds the question, searches a vector database, builds a prompt from the top results, calls the model, and post-processes the reply. An agent may loop through several model calls and tool calls before answering. When the final answer is wrong, the fault could sit in any of those steps, and a flat log does not show which step fed which.

Ordinary application-performance monitoring tools record timing and errors well, but they do not understand tokens, prompts, model names or the price of a call, and they have no concept of "this answer was 0.3 on faithfulness". Teams bolted those on by hand, differently every time.

The spreadsheet of test questions is the other half. People kept a column of inputs, pasted in the outputs by hand after each change, and eyeballed the differences. It is slow, it does not scale past twenty rows, and nobody can reproduce last week's run because nobody saved which prompt and model produced it.

Opik replaces the flat log with a tree of recorded steps and the spreadsheet with datasets and experiments that save the inputs, the outputs, the scores and the configuration of every run. Neighbouring tools solve parts of the same problem. If you later meet Langfuse, it is another open-source observability and evaluation platform with the same basic ideas (see the Langfuse guide), and RAGAS is a library of evaluation metrics for retrieval-augmented systems (see the RAGAS guide). Opik covers tracing, datasets, metrics, prompts and monitoring in one place, which is why it makes a good first tool to learn.

Why the vocabulary matters Opik's documentation uses a small set of words precisely: trace, span, thread, dataset, experiment, project. Learn what each one means in the next sections and every page of the interface becomes readable. Most beginner confusion comes from using "trace" to mean "anything logged", which is not what the word means here.
Try it
  1. Open the official documentation home page, https://www.comet.com/docs/opik/, and scan the left-hand navigation.
  2. Find the headings for Tracing, Evaluation, Prompt management and Production monitoring.
four groups of documentation that map onto the observe, evaluate and improve story above. You will touch the first two in this guide.

The mental model: workspace, project, trace, span

Opik has a handful of nouns, and they nest. Learn the nesting once and the interface stops being a maze.

WORKSPACE
tenant boundary: members, keys, limits
contains
PROJECT
one agent or app
contains
TRACE
one interaction
contains
SPANS
steps, as a tree

A workspace is the outermost boundary you will deal with. On Opik Cloud it holds your members, your AI-provider keys and your rate limits. On a self-hosted open-source install there is a single workspace called default and no user management at all, which matters for security and is covered in the Senior level.

A project represents one agent or application. In Opik 2.x, a project scopes not only the records of what your application did but also the datasets, experiments, prompts, online scoring rules, alerts and dashboards that belong to it. If you run a customer-support bot and a document summariser, you create two projects, and each keeps its own tests and prompts. If you do not choose a project, everything lands in one literally named Default Project.

That project-scoping is the biggest change in Opik 2.0 (released in spring 2026). In the 1.x series, only traces lived in projects, while datasets, prompts and experiments were shared across the whole workspace. Old blog posts and videos describe the 1.x behaviour. If an example you find online creates a dataset without a project_name and you cannot see it where you expect, this is why.

A trace is the complete record of one interaction with your application, one request and its response. It has a unique identifier, an input, an output, start and end times, optional metadata and tags, token usage and cost, and feedback scores. The guidance from the documentation is clear and worth memorising: a trace is one user interaction or one business operation. It is not a single function call, and it is not an entire session.

A span is one operation inside a trace. A trace for a RAG request might contain a span for retrieving documents, a span for building the prompt and a span for calling the model. Spans nest: a span can contain child spans, so the whole trace is a tree. Each span has a type. The common values are general (any ordinary step), llm (a call to a model), tool (a tool or function an agent uses) and guardrail. The type matters because only spans of type llm that carry a model name, a provider and token usage get a cost computed.

Here is what a trace for a tiny question-answering app looks like conceptually:

TEXT
Trace: answer_question            input: "What is our refund window?"   output: "30 days."
├── Span: retrieve_docs           type: general   (searches the knowledge base)
├── Span: build_prompt            type: general   (assembles instructions + context)
└── Span: call_model              type: llm       model: gpt-4o-mini   tokens: 412   cost: $0.0003

Reading a trace from the top tells you what the user saw. Reading down the tree tells you how that result was made, and the duration beside each span tells you where the time went. A slow answer stops being a mystery the first time you see that retrieve_docs took four of the five seconds.

Trace or span? Ask "would a user call this one request?" If yes, it is a trace. If it is a step that only makes sense as part of a request, it is a span. When in doubt, err toward fewer traces and more spans: a trace per function call floods the interface and hides the structure.
Try it
  1. Pick a small program you have written that has three or four functions calling each other.
  2. On paper, draw it as a tree: the entry-point function at the top, the functions it calls beneath.
  3. Decide which of the functions would be a trace and which would be spans.
exactly one trace at the top, with every other function as a span underneath. If you wanted several traces, your program probably handles several separate requests.

Threads, datasets, experiments and prompts

Four more nouns complete the picture. You will use all of them by the end of this guide, so a first pass now makes the later sections easier.

A thread groups traces that belong to one conversation. Each message a user sends to a chatbot is its own trace, and you tie them together by giving them a shared thread_id string that you choose. The identifier must be unique within the project. Opik then shows the conversation as a whole, and scoring rules can judge the conversation rather than a single reply. A thread is considered inactive after a cooldown (fifteen minutes by default), after which thread-level scoring can run.

A dataset is a collection of examples you test against. Each item has an input, usually an expected_output, and any other fields you want, such as a category. Datasets in Opik are versioned: every change produces a new immutable version, v1, v2 and so on, and a latest tag points at the newest. Inserting the same item twice does not duplicate it, because Opik deduplicates on insert.

An experiment is the record of one evaluation run. Opik takes each item in a dataset, hands it to your task (the function that runs your application), collects the output, scores it with one or more metrics, and stores everything together: every item, every output, every score and a trace for every item. You can open two experiments side by side to see whether a change helped. One rule is easy to forget: an experiment and all its traces land in the project of the dataset it ran against.

A metric is a scoring function. Heuristic metrics are deterministic and cheap: Equals checks for an exact match, Contains looks for a substring, RegexMatch applies a pattern, IsJson checks validity. LLM-as-a-judge metrics, such as Hallucination and AnswerRelevance, ask another model to grade the output and explain why. Every metric returns a score result with a name, a value and usually a reason.

A prompt in Opik's Prompt Library is a named, versioned template with {{variable}} placeholders. Changing the text creates a new version, and old versions stay available, so you can roll back and see which wording produced which results.

🔍

Trace

One request through your app, stored as a tree of spans.

💬

Thread

Several traces that share a conversation id.

📚

Dataset

Versioned examples to test against.

🧪

Experiment

One scored run of your task over a dataset.

There is also a newer object in 2.x, the test suite. It is a dataset-like collection where each item carries plain-English assertions ("the answer mentions the 30-day window") that an LLM judge checks. Test suites are covered at the Mid level; for now it is enough to know that they exist and that they suit behaviour you cannot express as an exact match.

Try it
  1. Write three questions and the answers you would accept for a tiny assistant, for example a bot that answers questions about your own university.
  2. Label each as input or expected_output.
a three-row table, which is a dataset. You will type it into Opik in a later section.

Choosing where Opik runs: Cloud or your laptop

Before installing anything, decide where the server lives. The SDK is only a client. It collects what your code does and sends it over HTTP to a server, and the server stores it and draws the web pages. So there are two separate things to set up: the SDK in your Python environment, and a server it can reach.

Opik Cloud

  • Nothing to run or upgrade
  • Free tier to start
  • Needs an API key and a workspace name
  • Data is stored on Comet's infrastructure
  • Best for learning and small teams

Local (Docker Compose)

  • Runs on your machine, data stays local
  • No account, no API key
  • Needs Docker and several gigabytes of memory
  • Opens at http://localhost:5173
  • Meant for learning, not production

For a first encounter, Cloud is the quickest: sign up at comet.com, create a workspace if prompted, and copy your API key from the user menu at the top right. The local route is better when you are offline, when you want to see what the moving parts are, or when your data must not leave your machine. Both routes are shown below and the rest of the guide works identically on either.

Two things in the local setup surprise newcomers. First, the open-source self-hosted version has no user accounts and no API keys: anyone who can reach the address can use it, so never expose a local instance to the internet. Second, a full local install runs several services (the application backend, a Python backend, the web front end, MySQL, ClickHouse, Redis, ZooKeeper and MinIO object storage), and ClickHouse in particular is hungry. Give Docker Desktop a few gigabytes of memory, or the stack will restart itself in the background.

If you have worked through the Docker guide, the local setup is a good exercise: you will see a real multi-container application start from a single script.

Try it
  1. Decide which route you will use for this guide. If you are unsure, pick Cloud.
  2. If Cloud: create the account now and copy your API key to a safe place.
  3. If local: check that docker --version and git --version both work.
either an API key in hand or two version numbers printed. Nothing else is needed before installation.

Installing the SDK and checking it works

The Python SDK is the most complete Opik client and the one this guide uses. It needs Python 3.10 or newer. Always install into a virtual environment so that Opik and its dependencies do not collide with other projects.

BASH
python3 -m venv .venv
source .venv/bin/activate        # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install -U opik
opik --version

The last command prints the installed version. You should see something in the 2.2 series. If you see 1.x, you are on an old release and most of this guide will not match, so upgrade with pip install -U opik.

Installing the package gives you two things: the importable library (import opik) and a command-line tool, also called opik, which includes the configure and healthcheck commands you will use in a moment.

Pointing the SDK at a server

The SDK needs to know where the server is and, for Cloud, who you are. The friendliest way is the interactive configuration command.

BASH
opik configure

For Cloud it asks for your API key, your workspace and a project name, then writes the answers to a file called ~/.opik.config in your home directory (on Windows, %USERPROFILE%\.opik.config). The project name you give becomes the default for everything you log. Because projects scope your datasets and prompts in Opik 2.x, choosing a meaningful name here, such as support-bot, rather than accepting the default is a good habit.

For a local server, use the --use_local flag, which writes the local address instead of asking for a key.

BASH
opik configure --use_local

The resulting file is plain text and looks like this (the exact keys present depend on your answers):

~/.opik.config
[opik]
url_override = http://localhost:5173/api
workspace = default
project_name = support-bot

For Cloud, the file instead holds api_key and workspace, and the address is the default Cloud URL, https://www.comet.com/opik/api, so no url_override is needed. The four keys you will meet most often are url_override, api_key, workspace and project_name.

You can do the same thing from Python, which is handy in notebooks:

PYTHON
import opik

opik.configure(use_local=False)   # Cloud: prompts for your key
# opik.configure(use_local=True)  # local server

Environment variables as an alternative

Every key has an environment-variable equivalent. Environment variables win over the file, which makes them the right tool for scripts, containers and continuous integration, where you should not commit a config file.

BASH
export OPIK_API_KEY="your-key-here"
export OPIK_WORKSPACE="your-workspace"
export OPIK_PROJECT_NAME="support-bot"
# for a local server instead:
export OPIK_URL_OVERRIDE="http://localhost:5173/api"

On Windows PowerShell the same thing is written $env:OPIK_API_KEY="your-key-here". The precedence, from strongest to weakest, is: arguments you pass in code, then environment variables, then ~/.opik.config, then built-in defaults.

Do not commit your key ~/.opik.config lives in your home directory, so it is safe by default. A .env file inside your project is not. Add .env to .gitignore before you create it, and never paste an API key into a notebook you will share.

Checking the connection

BASH
opik healthcheck

This command inspects your configuration and tries to reach the backend, and it is the first thing to run whenever anything misbehaves. A healthy result reports the URL it is using and that the connection succeeded. If it reports a problem, read the URL first. Nine out of ten beginner failures are a wrong address or a missing key, and the error section later in this guide lists the exact messages.

Try it
  1. Create a virtual environment and install opik.
  2. Run opik configure (or opik configure --use_local after starting the local server in the next section).
  3. Run opik healthcheck and read the output line by line.
  4. Open ~/.opik.config in an editor and find each key you just set.
a healthcheck that reports a reachable backend, and a config file whose values match what you typed. If it fails, note the message and jump to the errors section.

Running Opik locally with Docker

Skip this section if you chose Cloud. If you want your own server, the project ships a launcher script that starts the whole stack with Docker Compose. You need git, and Docker with Compose (Docker Desktop provides both on macOS and Windows).

BASH
git clone https://github.com/comet-ml/opik.git
cd opik
./opik.sh

On Windows, run the PowerShell version instead:

POWERSHELL
git clone https://github.com/comet-ml/opik.git
cd opik
powershell -ExecutionPolicy ByPass -c ".\opik.ps1"

The first run downloads several images, so expect a few minutes. When it finishes, the interface is at http://localhost:5173. Open it in a browser and you will see the Opik home page with no projects yet.

The script accepts flags that are worth knowing now, because you will use two of them constantly.

Flag What it does
--verify Checks that the containers are healthy
--stop Stops the stack and keeps your data
--clean Stops the stack and deletes all data volumes. Irreversible
--infra Starts only the databases and storage
--backend Starts infrastructure plus the backend
--guardrails Adds the optional guardrails backend
--build Builds images from source
--help Lists every option

Your data lives in a folder called opik inside your home directory, so stopping and restarting the stack keeps your traces. To upgrade, run ./opik.sh again. To pin a specific release, set the version first: OPIK_VERSION=2.2.86 ./opik.sh.

--clean really does delete everything It is the equivalent of docker compose down --volumes. Every trace, dataset and experiment you have recorded disappears. Use --stop when you only want to free up memory.

Once the stack is up, connect the SDK to it with opik configure --use_local, then check with opik healthcheck. One point trips up people who use TypeScript: the TypeScript SDK's default address is already the local server (http://localhost:5173/api), while the Python SDK's default is Cloud. The same code therefore behaves differently in the two languages until you configure it explicitly.

This local stack is meant for learning and development. It is not a production deployment. For production, Opik provides a Helm chart for Kubernetes, which is the subject of the Senior level, and connects naturally to the Helm guide and the Kubernetes guide.

Try it
  1. Run ./opik.sh and wait for it to finish.
  2. Run ./opik.sh --verify and read the health report.
  3. Open http://localhost:5173 in your browser.
  4. Run docker ps and count the containers.
a working web page and a half-dozen or more running containers, one per service in the architecture list above.

Your first traced function

Everything in Opik starts with one decorator. A decorator is a Python feature that wraps a function so that extra behaviour runs around it; you write @name on the line above the function definition. Opik's decorator is track, and it turns each call of the wrapped function into a recorded trace (or a span, if it is called from inside another tracked function).

Create a file called first_trace.py. This example calls no model, so it needs no provider key. That is deliberate: it lets you confirm Opik works before any other variable enters the picture.

first_trace.py
import opik
from opik import track


@track
def shout(text: str) -> str:
    return text.upper() + "!"


result = shout("hello opik")
print(result)

opik.Opik().flush()

Run it with python first_trace.py. You should see HELLO OPIK! printed, and a few seconds later a new project will appear in the web interface, named after the project you configured (or Default Project). Open it, go to the Logs page, and you will find one trace named shout. Click it. The input panel shows {"text": "hello opik"}, and the output panel shows the returned string. You never told Opik what the input and output were: @track read the function's arguments and return value for you.

The last line deserves an explanation because it is the most common reason beginners see nothing. Opik does not send each record the instant it happens. To keep your application fast, the SDK batches records in memory and sends them from a background thread. In a long-running server that is invisible. In a short script, the program can reach its last line and exit before the background thread has sent anything, and the trace is lost. Calling flush() tells the SDK to send everything that is queued and wait until it is done. You can also flush inside the decorator with @track(flush=True), which is convenient for small scripts and notebooks.

If you want to check quickly that the connection works without writing a file, the documented smoke test is this five-liner, which you can paste in a Python shell:

PYTHON
from opik import track
import opik

@track
def f(x):
    return x

f("hi")
opik.Opik().flush()
Silence is the usual failure When a trace does not appear, the cause is almost never that tracing is broken. It is one of four things: the script exited before flushing, the SDK is pointed at a different server or workspace than the page you are looking at, you are viewing a different project, or the API key is wrong. Check them in that order.
Try it
  1. Save the file above and run it.
  2. Find the trace in the Logs page of your project.
  3. Delete the flush() line, run the script again and see whether the new trace appears.
  4. Put the line back, then try @track(flush=True) instead.
the first run shows a trace. The second might or might not, because it is a race between your script exiting and the background sender. That uncertainty is the lesson: always flush in scripts.

Nested spans: seeing the tree

A single traced function is a trace with no children. The power appears when tracked functions call each other. The rule is simple: the outermost tracked call becomes the trace, and every tracked call made inside it becomes a span nested under its caller. You never wire the parent-child links yourself.

nested.py
import opik
from opik import track


@track(name="load-policy")
def load_policy(topic: str) -> str:
    policies = {"refund": "Refunds are accepted within 30 days of purchase."}
    return policies.get(topic, "No policy found.")


@track(name="write-answer")
def write_answer(question: str, policy: str) -> str:
    return f"Q: {question} | A: {policy}"


@track(name="answer-question")
def answer_question(question: str) -> str:
    policy = load_policy("refund")
    return write_answer(question, policy)


print(answer_question("How long do I have to return an item?"))
opik.Opik().flush()

In the interface, the Logs page shows one trace named answer-question. Open it and you see the tree in the left panel: answer-question at the root, load-policy and write-answer as its children. Clicking a span shows that span's own input and output, and each has a duration. This is the same structure as the conceptual diagram earlier, now produced by two lines of decoration.

Notice the name= argument. By default the span name is the function name, which is fine, but explicit names let you rename a function later without breaking the way you search and filter old traces.

The decorator accepts more options. The ones you will use in this first month are:

Option Purpose
name="..." The display name of the trace or span
type="llm" Marks the span as a model call, which enables cost tracking (other values: general, tool, guardrail)
project_name="..." Sends this trace to a specific project
flush=True Sends immediately after the function returns
environment="production" Tags the trace with a lifecycle label

If a tracked function raises an exception, Opik records the failure on the span and then lets the exception continue as normal. A failing step therefore shows up in the tree exactly where it failed, which is one of the most useful things tracing gives you.

Try it
  1. Run nested.py and open the trace.
  2. Add a third tracked function that is called from load_policy, then run again.
  3. Make write_answer raise ValueError("boom") and run once more.
a deeper tree after the second run, and after the third run a trace where the failing span is marked with an error and shows the exception. Finding that error takes one click, not a log search.

Tracing a real model call

Now add an actual model. You have two routes: wrap the provider's client so every call is traced automatically, or mark your own function as a model call by hand. Start with the wrapper, because it gives you model name, token counts and cost with no extra effort.

Opik has integrations for dozens of providers and frameworks, including OpenAI, Anthropic, Amazon Bedrock, Gemini, Ollama, LiteLLM, LangChain, LangGraph, LlamaIndex, CrewAI and the OpenAI Agents SDK. The pattern is nearly always the same: import a helper from opik.integrations, wrap your client or attach a callback, and carry on. Here it is for OpenAI.

BASH
pip install openai
export OPENAI_API_KEY="sk-..."
llm_call.py
import opik
from openai import OpenAI
from opik import track
from opik.integrations.openai import track_openai

client = track_openai(OpenAI())


@track(name="ask-model")
def ask(question: str) -> str:
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "Answer in one short sentence."},
            {"role": "user", "content": question},
        ],
    )
    return response.choices[0].message.content


print(ask("What does an LLM trace record?"))
opik.Opik().flush()

The wrapped client behaves exactly like the normal one. The only difference is that each chat.completions.create call now emits a span of type llm, nested under ask-model because ask is tracked. Open the trace and click the inner span. You will see the full list of messages sent, the model's reply, the model name, prompt and completion token counts and an estimated cost in dollars. The cost comes from a built-in price table keyed on the model, so a very new or custom model may show no cost rather than a wrong one.

Two practical points. First, track_openai takes an optional provider= argument if you point the OpenAI client at a compatible service, such as a gateway or a local server, so the span records the right provider. Second, you do not have to use the decorator at all: a wrapped client on its own produces a trace per call, but the decorator is what lets you group several calls into one meaningful trace.

If you use LangChain or LangGraph, the integration is a callback object. For LangGraph the current pattern is to wrap the compiled graph:

PYTHON
from opik.integrations.langchain import OpikTracer, track_langgraph

app = track_langgraph(app, OpikTracer())

An older pattern, where you pass OpikTracer(graph=app.get_graph(xray=True)) and add config={"callbacks": [tracer]} to each call, is still common in tutorials and still works, but the wrapper form is simpler. If you are learning LangGraph itself, see the LangGraph guide.

Doing it by hand when no integration fits

Sometimes you call a model through a library Opik has no wrapper for. You can still get full cost tracking by telling Opik what the span is. Put type="llm" on the decorator and update the span with the model, provider and usage.

manual_llm.py
import opik
from opik import track, opik_context


@track(name="custom-model-call", type="llm")
def call_my_model(prompt: str) -> str:
    answer = "stub answer"          # replace with a real call
    opik_context.update_current_span(
        model="my-model",
        provider="custom",
        usage={"prompt_tokens": 20, "completion_tokens": 5, "total_tokens": 25},
    )
    return answer


call_my_model("Hello")
opik.Opik().flush()

Cost is only calculated when the span is an llm span and carries a model, a provider and usage. Miss one of the three and the cost column stays empty. If the model is not in Opik's price table, cost is reported as missing, which is honest and safer than a guess.

Keys stay out of traces Opik records what you pass as function arguments and what you return. If you write a tracked function that takes your API key as a parameter, the key becomes part of the trace. Read secrets from the environment inside the function instead of passing them around.
Try it
  1. Run llm_call.py with your own key and find the trace.
  2. Click the inner LLM span and read the model, token counts and cost.
  3. Change the question and the model name, run again, and compare the two traces.
two traces whose LLM spans differ in tokens, latency and cost. If you have no provider key, run manual_llm.py instead and confirm the usage numbers you supplied appear on the span.

Adding context: tags, metadata, threads and feedback

A trace that only contains input and output is useful. A trace you can search, group and score is far more useful. Opik gives you four ways to attach extra information, and all of them are set from inside a tracked function through opik_context.

Tags are short labels, used for filtering: "support", "beta-user", "experiment-b". Metadata is a dictionary of arbitrary details: the model version, a customer tier, the git commit, the name of the retriever. Use tags for coarse categories you will filter by constantly, and metadata for everything else. A thread id groups traces into a conversation. Feedback scores are named numbers, optionally with a reason, recording a judgement: did the user like it, was it correct, was the tone right.

context.py
import opik
from opik import track, opik_context


@track(name="chat-turn")
def chat_turn(user_id: str, conversation_id: str, message: str) -> str:
    reply = f"You said: {message}"

    opik_context.update_current_trace(
        tags=["support", "demo"],
        metadata={"user_tier": "free", "user_id": user_id},
        thread_id=conversation_id,
        feedback_scores=[{"name": "user_feedback", "value": 1, "reason": "thumbs up"}],
    )
    return reply


chat_turn("u-42", "conv-001", "Hi there")
chat_turn("u-42", "conv-001", "What are your hours?")
opik.Opik().flush()

After this runs, open the project. The two calls are two separate traces, but they share the thread id conv-001, so the Threads view shows them as one two-message conversation. The tags appear as coloured chips you can filter on, the metadata is listed in the trace details, and user_feedback shows as a score column. You can filter the Logs page by any of those, for example to see only traces tagged support whose user_feedback is below one.

There is an important subtlety in the function names. update_current_trace changes the trace at the top of the current tree, while update_current_span changes only the span you are currently inside. When you want to record which retriever a particular step used, that belongs on the span. When you want to record which customer the whole request belonged to, that belongs on the trace.

PYTHON
opik_context.update_current_span(metadata={"retriever": "bm25", "top_k": 5})

The decorator can also receive this context at call time, without changing the function body, through an opik_args dictionary. That is useful when a framework calls your function for you and you cannot edit the inside.

Scores can also be added in the interface. Open a trace, find the feedback-scores area and add a score by hand, which is how human reviewers mark answers "good" or "bad". Names for human scores can be declared ahead of time as feedback definitions, so everyone uses the same vocabulary and the same scale. Later, the same score names will be written automatically by your evaluation metrics and by online scoring rules, so a single column can mix human and automatic judgement.

Threads need your help Opik cannot tell which traces belong to the same conversation. You must pass the same thread_id on every turn, and it must be unique per conversation within the project. Use your own session or conversation identifier rather than a random value generated on each call, or every message becomes a one-message thread.
Try it
  1. Run context.py.
  2. In the project, open the Threads view and find conv-001.
  3. Filter the Logs page by the tag support.
  4. Add a manual feedback score to one trace in the interface.
one thread containing two traces, a tag filter that keeps both, and a third score value you added by hand next to the two the code produced.

Projects and environments

Since Opik 2.x makes the project the home of everything an application owns, it pays to understand exactly how the SDK decides which project a trace goes to. The answer is an ordered list, and the first match wins:

  1. An explicit argument, such as @track(project_name="support-bot") or opik.Opik(project_name="support-bot").
  2. The environment variable OPIK_PROJECT_NAME.
  3. The project_name line in ~/.opik.config.
  4. The literal string Default Project.

If nothing is configured, your traces land in Default Project and the SDK logs a one-time warning. That is the usual reason people "lose" traces: they are there, just in a different project than the one they are viewing. Decide your project at the top of the program and keep it there.

projects.py
import opik
from opik import track

client = opik.Opik(project_name="support-bot")


@track(project_name="support-bot")
def handle(question: str) -> str:
    return question[::-1]


handle("hello")
client.flush()

There is one rule that surprises people. If you wrap code in opik.project_context("name"), the outermost project context wins. A nested @track(project_name="other") inside it is ignored, with a warning, by design, so that one trace never gets split across projects. Set the project at the outermost level and do not rely on inner overrides.

Environments are a second, lighter label. Where a project separates applications, an environment separates the stages of one application: development, staging, production. You can set it per call with @track(environment="production"), or globally with the OPIK_ENVIRONMENT variable, and filter the Logs page by it. Keeping staging and production in the same project, separated by an environment label, makes comparisons easy. Keeping them in separate projects gives cleaner isolation. Either is reasonable; what matters is to choose one and be consistent.

Try it
  1. Run projects.py with OPIK_PROJECT_NAME set to a different name.
  2. Check which project the trace landed in, then remove the project_name argument and run again.
the first run follows the explicit argument and ignores the environment variable, the second follows the variable. That is the precedence list in action.

Datasets: your test questions, saved

Tracing tells you what happened. Evaluation asks whether it was good. That needs a fixed set of examples so that you can run the same questions before and after every change, and in Opik that set is a dataset.

You create and fetch datasets through the client. get_or_create_dataset is the safe form: it returns the dataset if it exists and creates it if not, so the script can be run repeatedly.

make_dataset.py
import opik

client = opik.Opik(project_name="support-bot")

dataset = client.get_or_create_dataset(name="refund-questions", project_name="support-bot")

dataset.insert([
    {"input": "How long do I have to return an item?", "expected_output": "30 days"},
    {"input": "Can I get a refund after 45 days?", "expected_output": "No"},
    {"input": "Do refunds go to the original payment method?", "expected_output": "Yes"},
])

print("items:", len(dataset.get_items()))

Each item is a plain dictionary. Only the field names you choose matter, and later your task and metrics refer to them by name. Keep input and expected_output as conventions, since many built-in metrics look for those names, but nothing stops you adding category, difficulty or a source_document.

Run the script twice. The second run does not duplicate the items, because Opik deduplicates on insert. Each change produces a new immutable version of the dataset, and the interface lets you browse versions, so you can always tell which data an old experiment used. In the interface you can also add, edit and delete items in a draft, and committing the draft creates a new version.

Datasets can be loaded from files too. The client offers helpers such as insert_from_pandas and read_jsonl_from_file, which matter once your examples live in a spreadsheet or a JSONL export rather than in code.

Start with ten good rows A dataset of ten carefully chosen examples beats one of a thousand auto-generated ones. Include the easy cases, two or three awkward ones, and at least one the application should refuse or answer "I don't know". You can grow it whenever a real failure shows up in a trace: copy that failure into the dataset so it can never silently regress.
Try it
  1. Run make_dataset.py, then open the project's Datasets page.
  2. Open refund-questions and look at the items and the version list.
  3. Run the script a second time and confirm the item count does not change.
three items, a version number, and an unchanged count after the repeat run. Deduplication is why the script is safe to run again.

Running your first experiment

An experiment ties the pieces together: a dataset, a task that produces an answer for each item, and metrics that score each answer. The function that does this is evaluate.

The task is an ordinary Python function. Opik calls it once per dataset item, passing the item as a dictionary, and expects a dictionary back. The key names in the returned dictionary are what your metrics will read.

The example below uses a stand-in model so you can run it without a provider key. Swap in a real call later.

run_experiment.py
import opik
from opik.evaluation import evaluate
from opik.evaluation.metrics import Contains

client = opik.Opik(project_name="support-bot")
dataset = client.get_or_create_dataset(name="refund-questions", project_name="support-bot")


def task(item: dict) -> dict:
    question = item["input"]
    # Stand-in for a model call. Replace with your real application.
    if "return an item" in question:
        answer = "You have 30 days to return an item."
    elif "45 days" in question:
        answer = "No, that is outside the window."
    else:
        answer = "Yes, to the original payment method."
    return {"output": answer}


result = evaluate(
    dataset=dataset,
    task=task,
    scoring_metrics=[Contains(name="contains_expected")],
    scoring_key_mapping={"reference": "expected_output"},
    experiment_name="baseline-v1",
    project_name="support-bot",
    experiment_config={"model": "stand-in", "prompt_version": "none"},
)

When it finishes, Opik prints a summary and a link to the experiment. Under the hood, it ran your task on every item (using a pool of worker threads, eight by default), sent each output to each metric, and stored the results. Open the experiment page and you will see a table with one row per dataset item: the input, the expected output, what your task returned, and a score column for each metric.

The Contains metric checks whether a reference text appears inside the output. Its score method takes the task's output and a reference. Opik passes each metric the fields of the dataset item plus the task's returned dictionary, matched by name. Your dataset calls the field expected_output, but this metric asks for reference, so the scoring_key_mapping argument bridges the two: {"reference": "expected_output"} means "fill the metric's reference argument from the item's expected_output field". The same argument works for any renaming, for instance {"input": "user_question"} when your dataset calls the question field user_question. Mapping is the single most common thing beginners must learn about metrics, and a missing one shows up as an error about a missing argument.

The third question expects Yes and the stand-in returns Yes, ..., so it scores 1. The second expects No and the answer begins No,, so it also scores 1. Whether you care about case-sensitivity, exact match or a looser judgement is a choice you make by picking the metric, and this is the main design decision in evaluation.

Three other arguments from this call matter. experiment_name is the label you will search by later, so make it descriptive, for example baseline-v1 and then shorter-prompt-v2. experiment_config is a dictionary of whatever describes this run (model, prompt version, temperature), stored with the experiment so that a month from now you can tell what produced the numbers. project_name should match your dataset's project. Remember the rule from earlier: an experiment lands in the project of its dataset.

One limit to know early: evaluate does not accept an async task function. Write the task as an ordinary synchronous function; Opik provides the concurrency for you.

Try it
  1. Run run_experiment.py and open the experiment from the printed link.
  2. Change the stand-in answer for one question so it is wrong, then run again with experiment_name="baseline-v2".
  3. Open the Experiments page, select both runs and use the compare view.
two experiments whose average score differs, and a per-item comparison showing exactly which row got worse. That is regression testing for an LLM app.

Choosing metrics: heuristics and LLM judges

The metric decides what "good" means, so it deserves more thought than any other argument. Opik ships a large library, and they fall into three families.

Heuristic metrics are deterministic functions of the output. They are fast, free and perfectly repeatable, and they should be your first choice whenever the thing you care about can be checked by a rule. Equals demands an exact match. Contains looks for a substring. RegexMatch tests a pattern, useful for phone numbers or identifiers. IsJson checks that the output parses as JSON, which is invaluable for applications that must emit structured data. String-similarity metrics such as Levenshtein, BLEU, ROUGE, ChrF and BERTScore compare the output with a reference text when an exact match is too strict.

LLM-as-a-judge metrics ask another model to grade the output. They handle qualities no rule can capture: Hallucination asks whether the answer contains claims not supported by the provided context, AnswerRelevance asks whether it addresses the question, ContextPrecision and ContextRecall judge retrieval quality, Moderation flags unsafe content, and GEval lets you describe a custom criterion in plain English. Each returns a score between 0 and 1 and usually a written reason, which is where much of the debugging value lies.

Conversation metrics score a whole thread rather than one reply, for example whether a dialogue stayed coherent. They are covered at the Mid level.

judge_metric.py
from opik.evaluation.metrics import Hallucination

metric = Hallucination()
result = metric.score(
    input="How long is the refund window?",
    output="The refund window is 90 days.",
    context=["Refunds are accepted within 30 days of purchase."],
)
print(result.value, result.reason)

This needs a model to do the judging. By default the Python SDK uses openai/gpt-5-nano, configurable through the OPIK_DEFAULT_LLM environment variable, so an OPENAI_API_KEY must be set. To use a different judge, pass a model string in LiteLLM format, such as Hallucination(model="bedrock/..."). The judge reads your text, so anything in it leaves your machine for the provider. For regulated data, choose a provider and region that your employer permits, or a locally hosted model.

The output of a score is a ScoreResult with a name, a value and a reason. For the hallucination metric, a higher value means more hallucination, so read each metric's documentation before you interpret the direction. This is a frequent source of confused charts.

Use a heuristic when

  • The answer has a known right form
  • You need it free and repeatable
  • You run it on every commit
  • Checking format, presence or pattern

Use a judge when

  • Quality is a matter of meaning
  • You can afford provider tokens
  • You want a reason with each score
  • Checking faithfulness or relevance

A sensible starting combination is one or two cheap heuristics plus one judge. Use the heuristics to catch the broken cases cheaply, and the judge to watch the subtle quality. Judges are themselves imperfect and sometimes wrong, so read a sample of their reasons before you trust the averages. Adding a judge to the earlier experiment is a one-line change:

PYTHON
from opik.evaluation.metrics import Contains, Hallucination

scoring_metrics = [Contains(name="contains_expected"), Hallucination()]

Because the judge also needs the retrieved context, your task must return it under the key the metric expects, for example {"output": answer, "context": [doc1, doc2]}. When a metric complains about a missing argument, it is telling you exactly which key the task should return or the mapping should supply.

Writing your own metric When the built-in set is not enough you can subclass BaseMetric and implement score, returning a ScoreResult. Include **ignored_kwargs in the signature so the metric tolerates extra fields. Custom metrics are covered in depth at the Mid level.
Try it
  1. Add RegexMatch or IsJson to your experiment and decide what pattern would make sense for one of your questions.
  2. If you have a provider key, run the Hallucination snippet above and read the reason.
  3. Change the output so it agrees with the context and run it again.
a score that moves in the direction the documentation says it should, and a reason in plain sentences that tells you which claim the judge objected to.

Managing prompts in the Prompt Library

A prompt is code that happens to be written in English. It changes often, it affects behaviour, and a silent change can make quality drop. So it should be versioned like code. Opik's Prompt Library stores named, versioned prompt templates with {{variable}} placeholders, and lets your application fetch a specific version at runtime.

prompts.py
import opik

client = opik.Opik(project_name="support-bot")

prompt = client.create_prompt(
    name="refund-answerer",
    prompt="You are a support agent. Answer in one sentence.\nPolicy: {{policy}}\nQuestion: {{question}}",
    project_name="support-bot",
)

print(prompt.format(
    policy="Refunds are accepted within 30 days.",
    question="Can I return this?",
))

create_prompt is careful about versions. If a prompt with that name exists and the text is identical, nothing new is created. If the text changed, you get a new version, v2, v3 and so on; old versions are immutable. prompt.format(...) fills the placeholders and returns the final string, ready for a model call. To fetch an existing prompt, including a particular version:

PYTHON
latest = client.get_prompt(name="refund-answerer", project_name="support-bot")
pinned = client.get_prompt(name="refund-answerer", version="v1", project_name="support-bot")

There is a bonus that makes this worth doing from the start. When you call get_prompt inside a tracked function, Opik links that exact prompt version to the trace. Later, when one trace looks bad, you can see which wording produced it, and compare traces across versions. You can also pass prompts=[prompt] into evaluate so that an experiment records which prompt version it tested.

Prompts come in two kinds. A text prompt is a single string, as above. A chat prompt is a list of role-tagged messages (system, user, assistant) and uses create_chat_prompt and get_chat_prompt. Prompts are fetched and cached by the SDK, so edits made in the interface reach running code after a short delay (five minutes by default, controlled by OPIK_PROMPT_CACHE_TTL_SECONDS). That is helpful to know if an edit seems not to take effect.

Finally, prompts can be edited in the interface, and Opik's Playground lets you try a prompt against a model and even run it over a dataset, using provider keys you configure once per workspace under Configuration and AI Providers.

Try it
  1. Run prompts.py and find the prompt in the Prompt Library in the interface.
  2. Change one word in the template and run the script again.
  3. Open the prompt and look at the version list.
two versions, with the newer one marked as current. Running the script a third time with no change adds no version.

Reading the interface

The code you have written produces records, but you spend much of your time reading them, so it is worth learning the layout of the web interface. The names below match Opik 2.x.

The left navigation lists the areas of a project. Logs is the page you will use most: threads, traces and spans are all shown on this one page, with tabs to switch between them. Above the table are a search box, column controls and filters. Clicking a row opens a panel with the span tree on the left and the details on the right: input, output, metadata, feedback scores and usage. Experiments no longer have their own trace list elsewhere: the traces from an experiment appear as a Logs tab on the experiment page.

Dashboards and metrics show cost, token usage, latency and error counts over time, filtered by project. This is where the price of your application becomes visible. Datasets lists your collections and their versions. Experiments lists the runs, with a compare mode. Prompts is the library. Annotation queues let you assemble a batch of traces for human reviewers to score. Optimization runs (formerly "Optimization studio") manage automated prompt improvement. Rules hold online scoring, where a judge metric automatically scores a sample of your live traces.

Filtering uses a small language. In the filter bar, conditions look like <COLUMN> <OPERATOR> <VALUE>. Strings need double quotes, nested values use dots such as metadata.model or feedback_scores.user_feedback, and conditions are joined with AND. The language has no OR, which surprises people: to see either of two cases, run two filters.

There is also a built-in assistant called Ollie, which can read your traces and help you find patterns. It is a convenience layer, and it uses credits on Cloud. You do not need it to follow this guide.

One last interface trick is useful to know. Press Cmd+Shift+. on macOS, or Ctrl+Shift+. on Windows and Linux, to toggle Comet Debugger Mode, which displays the server version and the round-trip time to it. When a bug report is needed, or when you suspect the SDK and the server are at mismatched versions, this is the quickest way to read the server's version.

Try it
  1. On the Logs page, switch between the Threads, Traces and Spans tabs.
  2. Add a filter such as tags contains "support" or one on a feedback score.
  3. Find the cost and token charts on the metrics page.
the same data seen three ways, a filtered table that shrinks as you add conditions, and charts of the tokens and dollars your experiments used.

A short look at TypeScript and other languages

Everything so far used Python. If you write JavaScript or TypeScript, Opik has an SDK for Node, installed with npm install opik, and a setup command, npx opik-ts configure, which needs Node 18.17 or newer. Its concepts are identical, and its calls mirror the Python low-level client.

trace.ts
import { Opik } from "opik";

const client = new Opik({ projectName: "support-bot" });

const trace = client.trace({
  name: "answer-question",
  input: { question: "How long is the refund window?" },
  output: { answer: "30 days" },
});

trace.span({
  name: "call-model",
  type: "llm",
  input: { prompt: "..." },
  output: { text: "30 days" },
});

trace.end();
await client.flush();

There is no decorator here; instead you create the trace and its spans explicitly, end the trace and flush. Integrations ship as separate npm packages, for example opik-openai whose trackOpenAI wraps the OpenAI client and opik-vercel for the Vercel AI SDK.

For any other language, Opik offers a REST API and OpenTelemetry support: a standard way of exporting trace data that many languages already speak. Point an OpenTelemetry exporter at the Opik endpoint, https://www.comet.com/opik/api/v1/private/otel on Cloud, or http://localhost:5173/api/v1/private/otel locally, over HTTP (gRPC is not supported). Headers carry your key, project and workspace. The Mid level covers this properly.

Default addresses differ by language The TypeScript SDK assumes a local server unless told otherwise, while the Python SDK assumes Cloud. If the same settings work in one language and not the other, check the URL first.
Try it
  1. If you use Node, install opik, run npx opik-ts configure and run the snippet above.
  2. Otherwise, read the TypeScript SDK page on the official site and note where its names differ from Python.
a trace in the same project with the same shape as the Python ones. The interface cannot tell which language produced it.

Configuration you will actually touch

You have now met most of the configuration surface. This section gathers it in one place so you can look it up.

Setting Config file key Environment variable Notes
Server address url_override OPIK_URL_OVERRIDE Cloud default; local is http://localhost:5173/api
API key api_key OPIK_API_KEY Cloud only; self-hosted open source has none
Workspace workspace OPIK_WORKSPACE default on a self-hosted install
Project project_name OPIK_PROJECT_NAME Falls back to Default Project
Environment (none) OPIK_ENVIRONMENT Lifecycle label such as production
Switch tracing off opik_track_disable OPIK_TRACK_DISABLE Also opik.set_tracing_active(False)
Judge model (none) OPIK_DEFAULT_LLM Default openai/gpt-5-nano
Usage analytics analytics_enable OPIK_ANALYTICS_ENABLE On by default in the Python SDK; set to false to opt out

Some observations on that table. The config file is the friendliest place for values that belong to you personally and never change, and environment variables are better for anything a script or container needs to override. The OPIK_TRACK_DISABLE switch is worth knowing: in unit tests you usually do not want your code sending traces, so set it to true in the test environment, or call opik.set_tracing_active(False).

Older material sometimes uses OPIK_BASE_URL for the server address. The current configuration reference uses OPIK_URL_OVERRIDE, so use that one.

Logging deserves one warning. To see what the SDK itself is doing while debugging, use the environment variables OPIK_CONSOLE_LOGGING_LEVEL (default INFO) and OPIK_FILE_LOGGING_LEVEL, and set them before your program imports opik. The usual Python approach, logging.getLogger("opik").setLevel(...), has no effect, because the SDK manages its own logging. Likewise, if you load secrets from a .env file with python-dotenv, call load_dotenv() before import opik, or the values will be read too late and ignored.

PYTHON
from dotenv import load_dotenv
load_dotenv()          # first

import opik            # then import opik
Try it
  1. Run a traced script with OPIK_CONSOLE_LOGGING_LEVEL=DEBUG set and look at the extra output.
  2. Set OPIK_TRACK_DISABLE=true and run it again.
verbose lines describing the SDK's connection and batching in the first run, and no new trace at all in the second, even though the function still returns normally.

When things go wrong: common errors

Most of the trouble you will have falls into a handful of patterns. The format here is the message you will see or the symptom, what causes it, and the fix.

Nothing appears in the interface. The most frequent cause in short scripts is that the program exited before the background sender delivered the batch. Call opik.Opik().flush() at the end, or use @track(flush=True). If flushing does not help, run opik healthcheck and compare the URL, workspace and project it prints with the ones in your browser.

Traces appear in Default Project. You did not configure a project, or a project context overrode your intention. Set OPIK_PROJECT_NAME, pass project_name, or rerun opik configure. A one-time warning in the console says this happened.

HTTP 403 Forbidden. The API key is wrong or missing, you are pointing at a workspace you cannot access, or the URL is wrong. Re-run opik configure and check OPIK_API_KEY, OPIK_WORKSPACE and OPIK_URL_OVERRIDE. For a self-hosted server the URL must end in :5173/api; the /api is because the front end's web server forwards that path to the backend.

A certificate error.

TEXT
[SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: self-signed certificate in certificate chain (_ssl.c:1006)

This appears when your company network inspects HTTPS traffic with its own certificate, or when a server uses a self-signed one. The correct fix is to tell Python to trust your organisation's certificate authority by setting REQUESTS_CA_BUNDLE to the path of its certificate bundle. Setting OPIK_CHECK_TLS_CERTIFICATE=false also makes the error disappear, but it turns off verification, so treat it as a short-lived last resort rather than a solution.

HTTP 429 and a log line such as OPIK: Ingestion rate limited, retrying in 55 seconds. Opik Cloud limits how fast you may send. The SDK waits and retries automatically, so the message is informational. If you hit it constantly, reduce how many threads you run in parallel.

HTTP 413, payload too large. A single field was enormous, typically a full document or a base64 image inside an input or output. Log identifiers, summaries or the top few results instead of entire blobs.

Updates lost and a batching warning. If you use the lower-level client and call .end() or .update() on a trace immediately after creating it, the update can arrive before the create and be dropped. The decorator and context managers avoid the problem. This matters more at the Mid level.

A metric fails with an unexpected-argument or missing-argument error. The keys your task returns and the dataset fields do not match what the metric's score method expects. Print what the task returns, read the metric's documentation for its argument names, and add a scoring_key_mapping.

An evaluation refuses an async task. evaluate does not support async def tasks. Make the task a normal function.

Datasets, prompts or experiments seem to vanish after upgrading from 1.x on a self-hosted server. The 2.0 release made these project-scoped, and old items without a project are hidden until migration jobs run. On Cloud the migration happened automatically. On self-hosted the procedure is in the official migration guide and is a Senior-level topic, but the symptom is worth recognising.

uvx: command not found when following MCP instructions. That is the uv tool missing, or a terminal opened before installing it. Install uv, then open a new terminal.

A reliable debugging order First opik healthcheck. Then confirm the URL, workspace and project. Then check the script flushes. Then turn on OPIK_CONSOLE_LOGGING_LEVEL=DEBUG. Only after all four should you suspect a bug in Opik itself.
Try it
  1. Set OPIK_URL_OVERRIDE to a wrong address such as http://localhost:9999/api and run opik healthcheck.
  2. Read the error, then restore the correct value and run it again.
a connection failure that names the address it tried, and a clean pass once restored. Deliberately breaking the setup once is the fastest way to recognise the real thing later.

Putting it all together

Here is one small end-to-end project that uses everything above: a tiny question-answering assistant that is traced, has a versioned prompt, is tested against a dataset, and is scored. It has no external dependency beyond Opik so you can run it anywhere, and a marked spot where you can drop in a real model.

Create a folder called refund-bot and add these files. First, the application.

refund-bot/app.py
import opik
from opik import track, opik_context

PROJECT = "refund-bot"
client = opik.Opik(project_name=PROJECT)

POLICY = "Refunds are accepted within 30 days of purchase. Refunds go to the original payment method."

prompt = client.create_prompt(
    name="refund-answerer",
    prompt="Policy: {{policy}}\nQuestion: {{question}}\nAnswer in one sentence.",
    project_name=PROJECT,
)


@track(name="retrieve-policy")
def retrieve_policy(question: str) -> str:
    return POLICY


@track(name="generate", type="llm")
def generate(filled_prompt: str) -> str:
    # Replace this stub with a real model call, for example via track_openai.
    text = filled_prompt.lower()
    if "how long" in text or "return" in text:
        return "You have 30 days to return an item."
    if "45 days" in text:
        return "No, that is outside the 30 day window."
    return "Yes, refunds go to the original payment method."


@track(name="answer-question", project_name=PROJECT)
def answer_question(question: str, conversation_id: str = "demo") -> str:
    policy = retrieve_policy(question)
    filled = prompt.format(policy=policy, question=question)
    answer = generate(filled)
    opik_context.update_current_trace(
        tags=["refund-bot"], thread_id=conversation_id, metadata={"prompt_version": "latest"}
    )
    return answer

Then the evaluation script, which imports the application and scores it against a dataset.

refund-bot/evaluate.py
import opik
from opik.evaluation import evaluate
from opik.evaluation.metrics import Contains

from app import PROJECT, answer_question

client = opik.Opik(project_name=PROJECT)
dataset = client.get_or_create_dataset(name="refund-questions", project_name=PROJECT)
dataset.insert([
    {"input": "How long do I have to return an item?", "expected_output": "30 days"},
    {"input": "Can I get a refund after 45 days?", "expected_output": "No"},
    {"input": "Do refunds go to the original payment method?", "expected_output": "original payment method"},
])


def task(item: dict) -> dict:
    return {"output": answer_question(item["input"])}


evaluate(
    dataset=dataset,
    task=task,
    scoring_metrics=[Contains(name="contains_expected")],
    scoring_key_mapping={"reference": "expected_output"},
    experiment_name="refund-bot-baseline",
    project_name=PROJECT,
    experiment_config={"prompt": "refund-answerer", "model": "stub"},
)
client.flush()

Run it from inside the folder:

BASH
cd refund-bot
python evaluate.py

What you should see and what to do with it. The console prints an evaluation summary and a link. In the interface, the refund-bot project now contains one experiment, refund-bot-baseline, with three rows. Open any row's trace from the experiment's Logs tab and you will see the tree you built: answer-question at the top with retrieve-policy and generate beneath it, and the generate span marked as an LLM span. The Prompts page shows refund-answerer. The Datasets page shows refund-questions. Every object belongs to one project, as Opik 2.x intends.

Now make a change and watch it register. Edit the prompt text in app.py, or change the stub's answer for the first question to something wrong, and run python evaluate.py again with a new experiment_name, for example refund-bot-edit-1. Open the Experiments page, select both runs and compare. You have just done what evaluation-driven development means: propose a change, measure it, and decide from data.

Finally, swap the stub for a real model. Use track_openai as shown earlier, call the model inside generate, and mark it with type="llm". Add Hallucination() to scoring_metrics and return "context": [POLICY] from the task. Everything else stays the same, which is the point of the design: the harness does not care what is inside generate.

Try it
  1. Build the refund-bot folder and run python evaluate.py.
  2. Find the experiment, its trace tree, the prompt and the dataset in the interface.
  3. Break one answer on purpose, rerun under a new experiment name and compare the two runs.
  4. Add a fourth dataset item that the bot should refuse, and see how the bot fares.
a complete loop: code, traces, dataset, experiment, comparison. The fourth item will probably fail, and a failing example you wrote on purpose is how a test set starts to earn its keep.

What you can now do, and what comes next

If you have worked through the Try-it tasks, you can now do things that many teams skip. You can explain the difference between a workspace, a project, a trace and a span without hesitating. You can install the SDK, point it at Cloud or a local server, and verify the connection with opik healthcheck. You can trace any Python function with @track, nest functions to get a tree, wrap a model client to capture tokens and cost, and attach tags, metadata, thread ids and feedback scores. You can build a versioned dataset, run evaluate with a task and metrics, compare two experiments, and keep your prompts in the Prompt Library. And you know the usual reasons for silence: the missing flush, the wrong project, the wrong URL.

Keep these habits from day one. Set a project name explicitly, and never rely on Default Project. Flush in every script. Name experiments so that you can read the list a month later. Put the model, the prompt version and any other settings into experiment_config. Keep secrets out of function arguments. Add a failing real-world example to your dataset every time production surprises you.

The Mid-level guide starts where this one stops. It explains how batching and the client work precisely enough to predict their behaviour, covers the low-level client and distributed tracing across services, custom metrics and test suites with natural-language assertions, online evaluation of production traffic, annotation queues, the command-line export and import tools, OpenTelemetry, and integration into continuous delivery. The Senior guide treats Opik as a platform for a team: its architecture, the ClickHouse-centred scaling limits, security and access control, upgrade and migration procedures such as the 1.x to 2.x move, backup, cost control and incident playbooks.

Neighbouring tools are worth knowing as you grow. Langfuse solves an overlapping problem with a different philosophy, RAGAS provides retrieval-specific metrics, and LangGraph is a common framework whose runs you will want to trace. The MLflow guide shows how classical machine-learning experiment tracking compares.

Sources