Skip to content
Back to student guides
TruLensLLMOpsEvaluation & testing3 levels99 sectionsCovers TruLens 2.14

The Complete TruLens Guide

Evaluate and track LLM apps with TruLens feedback functions. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
13sections
20examples

This is part one of three. It covers everything you need to do real work with TruLens, not a teaser. By the end you can take a small question-answering app that calls a language model, record every run of it, score each answer for relevance and for hallucination, open a dashboard that compares two versions of the app, and read the scores well enough to decide which version to ship. Mid-level and Senior take the same topics further; nothing here is thrown away.

This guide is written against TruLens 2.14.0, the release current in September 2026. TruLens changes quickly, and a lot of the blog posts and tutorials you will find online describe older APIs. Where an older pattern is still common on the internet, this guide names it and shows the current form, so you can recognise a stale tutorial when you meet one.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas stick only after you have seen a score appear for an answer you produced yourself.

Why evaluating an LLM app is hard, and what TruLens does about it

Ordinary software has a comforting property: give it the same input and you get the same output, so you can write a test that says "this input must produce exactly that output". An application built on a large language model (LLM, a model that generates text) does not behave that way. Ask it the same question twice and the wording differs. Change one sentence in the prompt, swap the model for a cheaper one, or re-index your documents, and the answers shift in ways that no assert answer == "..." line can capture. Worse, the failures are not crashes. The app returns a fluent, confident, well-formatted paragraph that happens to be wrong, or that answers a different question, or that quietly invents a fact that appears in none of your documents. The word for that last failure is hallucination.

Teams used to handle this by reading outputs by hand. Someone pastes twenty questions into the app, eyeballs the answers, says "looks fine", and ships. That works exactly once. The second time you change something, you have no record of what the twenty answers looked like before, no way to tell whether the new version is better or merely different, and no way to read two hundred answers a day once real users arrive. You need two things that manual reading cannot give you: a record of what the app did on every call, and a score for each call that you can average, compare, and watch over time.

TruLens is an open-source Python library that provides exactly those two things. It does three jobs, and it helps to keep them apart in your head:

  1. Tracing. It watches your app run and records each step as a structured event: the question that came in, the documents that were retrieved, the prompt that was sent to the model, the answer that came back, how long each step took, and what it cost.
  2. Evaluation. It runs metrics over those recordings. A metric is a small function that looks at part of a recording and returns a score between 0 and 1. Many of TruLens's metrics use a second language model as a judge: the judge reads the question and the answer and says how relevant the answer is.
  3. Storage and display. It saves the recordings and scores in a database and gives you a dashboard, so that you can compare version A of your app with version B on the same numbers.

The loop that results is the heart of the tool, and the rest of this guide fills in each step of it:

INSTRUMENTmark the steps
→
RECORDrun the app
→
EVALUATEscore each run
→
COMPAREversions side by side
→
ITERATEchange, repeat

A little history explains some of the names you will meet. The project began at a company called TruEra, which Snowflake later acquired; Snowflake now maintains it, and the code is MIT-licensed and lives at github.com/truera/trulens. The early library was called trulens_eval, and its central object was called Tru. In version 1.0 (2024) the library was split into separate packages under a single trulens namespace, Tru became TruSession, and the old trulens_eval package became an empty shim that exists only to print a deprecation notice. In 2.3.0 (August 2025) the tracing machinery switched to OpenTelemetry by default, and in 2.7.0 (February 2026) the word "feedback function" gave way to "metric". You will still see all the old words in search results. The current words are the ones in this guide.

Do not copy code that imports trulens_eval If a tutorial starts with from trulens_eval import Tru, Feedback it was written for the pre-1.0 library and will not teach you the current API. The same goes for Tru().run_dashboard(). Everything in this guide uses the trulens.* namespace.

What TruLens is not matters too. It is not a framework for building the app; you build the app with plain Python, LangChain, LlamaIndex, or whatever you like, and TruLens watches it. It is not a hosted service: the core library runs inside your own Python process and writes to a file or database you control. And it is not a magic truth detector. Its scores come from other language models or from simple code, and they are evidence to reason about, not verdicts. A later section returns to that point, because treating a judge's score as ground truth is the most common conceptual mistake beginners make.

The places TruLens sits next to other tools are worth knowing. Langfuse and LangSmith are tracing and evaluation platforms with a similar purpose and different emphasis; RAGAS and DeepEval are libraries that focus on scoring. This catalogue has guides for them, and you can read them afterwards to compare how each approaches the same problem. TruLens's particular stance is that evaluation should be built from a small set of named, explainable metrics, and that every score should be traceable back to the exact spans of the exact run that produced it.

Try it
  1. Think of any LLM-backed feature you have used or built: a chatbot, a summariser, a search box with generated answers.
  2. Write down three ways its answer could be bad without the program crashing.
  3. For each, write how you would notice at a hundred answers a day.
a list of failures such as "answered a different question", "made up a number", "rambled". Each one is something a metric can score, which is the reason the rest of this guide exists.

The mental model: five nouns

TruLens has a long API, but a beginner needs only five nouns. Learn them now, because every error message and every doc page assumes you know them.

Session. The TruSession is the object that owns everything stateful: the connection to the database, the machinery that exports recordings, and the background workers that run metrics. There is only one per Python process. Calling TruSession() a second time returns the same object rather than creating a new one, which has a consequence we will meet in the errors section: you cannot change the database after the first call.

App. Your application, wrapped. You hand TruLens your Python object and give it two labels, an app_name and an app_version. The wrapper is called a recorder. For a plain Python class the wrapper is TruApp; for LangChain it is TruChain, for LangGraph TruGraph, for LlamaIndex TruLlama. The app_name plus app_version pair is what the dashboard groups by, so "RAG / v1" and "RAG / v2" are two entries you can compare. Choosing good version labels is a real skill, and the habits section of this guide and the tips file come back to it.

Record (and span). One run of your app is a record. Inside the record, each step is a span: one retrieval, one model call, one tool call. Spans nest, so the record as a whole is the outermost span, and the retrieval and generation sit inside it. Every span has a span type, which tells TruLens what kind of step it was. The types you need first are RECORD_ROOT (the whole run), RETRIEVAL (fetching documents), and GENERATION (asking the model). The word "record" is slated to be renamed "trace" in future releases, so you will see both.

Metric. A scoring function. You create one with Metric(...), and it needs three things: the function that does the scoring (the implementation), a selector for each of the function's inputs saying which part of the recording to feed it, and a name. Old tutorials call this a "feedback function" and use a class called Feedback. Feedback still works in 2.14.0, but constructing one prints a deprecation warning, and it will eventually be removed.

Provider. The model backend that does the judging. OpenAI is a provider; so are Anthropic, Google, Bedrock, LiteLLM, and a local one for Ollama. A provider object has many ready-made scoring methods, such as relevance and groundedness_measure_with_cot_reasons, and you pass one of those methods as the implementation of a metric.

Put together, the picture looks like this:

Your process
Your app
plain Python, instrumented
→
TruApp recorder
app_name + app_version
→
TruSession
spans out, metrics run
Judging
Metric
selectors pick the inputs
→
Provider
a judge model such as OpenAI
Storage and display
Database
SQLite file by default
→
Dashboard
a Streamlit web page

Two more ideas complete the model. The first is the selector. A metric such as "is this answer relevant to the question" needs two strings: the question and the answer. But your app does not hand TruLens a tidy pair; it records a tree of spans with attributes. A selector is a pointer into that tree: "the input of the record root", "the output of the record root", "every retrieved chunk". You attach one selector to each parameter of the metric's function, and the parameter names must match, which is the source of the most common beginner error. We will spend a whole section on it.

The second is the RAG Triad. RAG means retrieval-augmented generation: instead of asking the model a question cold, your app first searches a collection of your own documents, then pastes the best matches into the prompt, then asks the model to answer using them. It is the most common shape of LLM app in practice, and it can fail in three separate places. The retrieval can fetch the wrong documents. The model can ignore the documents and make something up. Or the model can answer faithfully from the documents but miss the question. TruLens's headline evaluation has one metric for each failure:

Metric Compares Catches
Context relevance the question with each retrieved chunk retrieval fetched the wrong documents
Groundedness the answer with the retrieved chunks the model made claims the documents do not support (hallucination)
Answer relevance the question with the answer the answer does not actually address the question

If all three are high, the app retrieved the right material, stuck to it, and answered what was asked. If one is low, it tells you which part to fix, and that diagnostic power is what makes the triad more useful than a single "quality" number.

Try it
  1. Imagine a company help bot asked "How many vacation days do I get?" and answering "Our office is open 9 to 5", after retrieving the office-hours page.
  2. Decide which of the three triad metrics would be low.
  3. Now imagine it retrieved the right policy page but said "You get 40 days". Which metric is low now?
in the first case context relevance and answer relevance are low; in the second, groundedness is low because 40 is not in the source. Being able to say which metric flags which failure is the core skill.

Installing TruLens and checking the setup

TruLens installs with pip and is split into several packages, so that you only download what you use. You need Python 3.10 to 3.13 for the main trulens package. A virtual environment is strongly recommended, because the library pulls in a fair number of dependencies and you do not want them fighting with your other projects.

On Linux or macOS:

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install trulens trulens-providers-openai

On Windows (PowerShell):

POWERSHELL
py -3.12 -m venv .venv
.venv\Scripts\Activate.ps1
pip install trulens trulens-providers-openai

The two packages do different things. trulens is a meta-package that installs the core library (trulens-core), the built-in metrics (trulens-feedback), the dashboard (trulens-dashboard), and the attribute definitions (trulens-otel-semconv). It deliberately does not install any provider or framework integration. trulens-providers-openai is the provider we will use as our judge. If you would rather judge with another model, install the matching package instead. The common ones are trulens-providers-anthropic, trulens-providers-google, trulens-providers-bedrock, trulens-providers-litellm (one package that reaches over a hundred models), and trulens-providers-ollama for models running on your own machine.

The wrappers for frameworks follow the same pattern. If your app is written with LangChain, add trulens-apps-langchain; for LangGraph, trulens-apps-langgraph; for LlamaIndex, trulens-apps-llamaindex. The wrapper for a plain Python class, TruApp, ships inside trulens-core and needs nothing extra.

The judge needs credentials. For OpenAI, set the key as an environment variable before you run anything. TruLens also reads a .env file in the current directory or a parent, so either approach works:

BASH
export OPENAI_API_KEY="sk-..."          # Linux / macOS
POWERSHELL
$env:OPENAI_API_KEY = "sk-..."          # Windows PowerShell

Never paste a key into a notebook cell or a file you will commit. If you use a .env file, add it to .gitignore first. Other providers read their own variables: ANTHROPIC_API_KEY for Anthropic, AZURE_OPENAI_API_KEY and AZURE_OPENAI_ENDPOINT for Azure OpenAI, and the usual cloud credentials for Bedrock and Google.

Now verify the install. The first command checks that the packages import and prints the version; the second lists everything TruLens-related that pip installed:

BASH
python -c "import trulens.core, trulens.dashboard; from importlib.metadata import version; print(version('trulens-core'))"
pip list | grep -i trulens

The first command should print 2.14.0, or a later version if you are reading this after a new release. On Windows, replace grep -i with findstr. One subtlety worth knowing: all the trulens-* packages are released together and share a version number. If pip list shows trulens-core at one version and a provider at a different one, upgrade them together, because mismatched versions are a classic source of confusing import errors.

Finally, create a session:

PYTHON
from trulens.core import TruSession

session = TruSession()

The first time, TruLens prints a line that begins with a squid emoji and says Initialized with db url sqlite:///default.sqlite. That line tells you where your data lives: a file called default.sqlite in whatever directory you ran Python from. Two beginner surprises follow from that. The first is that if you run the same script from a different directory, you get a different, empty database, and wonder where your results went. The second is that the file keeps growing as you run things; we will see how to reset it in a moment.

Run everything from one folder Make a project folder, create the virtual environment inside it, and always launch your scripts and the dashboard from that folder. Your default.sqlite will then be in the one place you expect.

If you are offline or do not want to spend money on a hosted judge while learning, you can run a model locally with Ollama and use the trulens-providers-ollama package. Small local models give noisier scores than large hosted ones, but the workflow is identical, and for learners in regions where sending prompts to a foreign API is a data-residency concern, a local judge is a legitimate default. The mid and senior guides return to that choice.

Try it
  1. Create a virtual environment and install trulens and trulens-providers-openai.
  2. Run the version command and confirm the number.
  3. Run the three-line session snippet and find the default.sqlite file it created.
the version prints, the squid line names your database, and the file exists in the folder you ran from. If an import fails, run pip list and check that trulens-core is present.

Your first project, step by step

We will build a tiny RAG app, record it, score it, and look at the results. To keep the moving parts few, the "document collection" is a Python list and the retriever is a keyword match; in a real project this is where a vector database goes, and the instrumentation you learn here does not change when you swap it. The generator calls OpenAI through its normal client, which the provider package installs for you.

Create a file called app.py. We build it in four pieces, then run it.

Step 1: the app, with no TruLens yet

app.py
from openai import OpenAI as OpenAIClient

DOCS = [
    "Employees receive 25 days of paid annual leave per calendar year.",
    "Unused leave can be carried over, up to a maximum of 5 days.",
    "The office is open Sunday to Thursday, from 9 am to 5 pm.",
    "Remote work is allowed up to two days per week with manager approval.",
]

client = OpenAIClient()


class HRBot:
    def retrieve(self, query: str) -> list[str]:
        words = {w.lower().strip("?.,") for w in query.split()}
        scored = [(len(words & set(d.lower().replace(".", "").split())), d) for d in DOCS]
        scored.sort(reverse=True)
        return [d for score, d in scored[:2] if score > 0]

    def generate_completion(self, query: str, context_list: list[str]) -> str:
        context = "\n".join(context_list)
        response = client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[
                {"role": "system", "content": "Answer using only the context. If it is not there, say you do not know."},
                {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {query}"},
            ],
        )
        return response.choices[0].message.content

    def query(self, query: str) -> str:
        context_list = self.retrieve(query)
        return self.generate_completion(query, context_list)

Read it as three steps: retrieve finds up to two documents that share words with the question, generate_completion asks the model to answer from them, and query is the public entry point that chains the two. This is ordinary Python. It is worth running HRBot().query("How many days of annual leave do I get?") once by itself, before adding any TruLens, so you know the app works and your key is valid. Debugging an instrumented app whose underlying app is broken is a miserable way to spend an afternoon.

Step 2: mark the steps with @instrument

TruLens cannot watch code it has not been told about. You tell it by adding the @instrument decorator to the methods that are steps, and you say what type of step each is and which values to capture. Add these imports at the top of the file and decorate the three methods:

app.py
from trulens.core.otel.instrument import instrument
from trulens.otel.semconv.trace import SpanAttributes


class HRBot:
    @instrument(
        span_type=SpanAttributes.SpanType.RETRIEVAL,
        attributes={
            SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
            SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
        },
    )
    def retrieve(self, query: str) -> list[str]:
        ...  # same body as before

    @instrument(span_type=SpanAttributes.SpanType.GENERATION)
    def generate_completion(self, query: str, context_list: list[str]) -> str:
        ...  # same body as before

    @instrument(
        span_type=SpanAttributes.SpanType.RECORD_ROOT,
        attributes={
            SpanAttributes.RECORD_ROOT.INPUT: "query",
            SpanAttributes.RECORD_ROOT.OUTPUT: "return",
        },
    )
    def query(self, query: str) -> str:
        ...  # same body as before

The ... lines stand for the bodies you already wrote; keep them. The attributes dictionaries are the important part. The keys are the names TruLens uses for well-known facts, from SpanAttributes. The values are strings that say where to find the fact in the function call: the name of one of its parameters, or the special word "return" for the returned value. So RETRIEVAL.QUERY_TEXT: "query" says "the retrieval's query text is the argument called query", and RETRIEVAL.RETRIEVED_CONTEXTS: "return" says "the retrieved contexts are whatever this method returns". On the record root, INPUT is the question that came in and OUTPUT is the final answer.

Why all this ceremony? Because the metrics need to find these values later. The groundedness metric needs "the retrieved contexts", and it locates them by looking for a span of type RETRIEVAL with the RETRIEVED_CONTEXTS attribute. If you forget to mark the retrieval, nothing crashes; the metric simply finds no context and produces no score. That is the quietest and most common failure in the whole tool, and this paragraph is the cure.

Step 3: define the metrics

Now create the judge and the three RAG Triad metrics. Put this below the class:

app.py
import numpy as np
from trulens.core import Metric, Selector, TruSession
from trulens.apps.app import TruApp
from trulens.providers.openai import OpenAI

session = TruSession()
session.reset_database()          # start from an empty database while learning

provider = OpenAI(model_engine="gpt-4o-mini")

f_groundedness = Metric(
    implementation=provider.groundedness_measure_with_cot_reasons_consider_answerability,
    name="Groundedness",
    selectors={
        "source": Selector.select_context(collect_list=True),
        "statement": Selector.select_record_output(),
        "question": Selector.select_record_input(),
    },
)

f_answer_relevance = Metric(
    implementation=provider.relevance_with_cot_reasons,
    name="Answer Relevance",
    selectors={
        "prompt": Selector.select_record_input(),
        "response": Selector.select_record_output(),
    },
)

f_context_relevance = Metric(
    implementation=provider.context_relevance_with_cot_reasons,
    name="Context Relevance",
    selectors={
        "question": Selector.select_record_input(),
        "context": Selector.select_context(collect_list=False),
    },
    agg=np.mean,
)

Take it slowly, because this block holds most of the vocabulary. provider = OpenAI(model_engine="gpt-4o-mini") creates the judge and pins its model. Always name the model explicitly; the class has a default, but defaults change between releases, and an unpinned judge means your scores can shift when you upgrade the library rather than when you change your app.

Each Metric takes an implementation, which here is a method of the provider, a name that appears in the dashboard, and a dictionary of selectors. The dictionary's keys must match the parameter names of the implementation. The relevance method is declared as relevance(prompt, response), so the keys are "prompt" and "response". Context relevance is context_relevance(question, context), and groundedness takes source and statement and optionally question. Get a key wrong and the metric fails or silently scores nothing. The names ending in _with_cot_reasons return a score plus a written explanation (CoT stands for chain of thought: the judge reasons step by step and TruLens keeps the reasoning), and that explanation is what you will read in the dashboard to understand a low score.

The selectors deserve a second look. Selector.select_record_input() and select_record_output() point at the question and answer captured on the record root. Selector.select_context(collect_list=True) points at the retrieved contexts. The collect_list flag decides how many judge calls happen. With True, all the chunks are passed together as one list and the judge scores them in one call; that is right for groundedness, where the question is "is the answer supported by the documents as a whole". With False, each chunk is scored separately and the scores are combined with the agg function, here np.mean; that is right for context relevance, where you want to know how relevant each chunk is on average. The default aggregation is also the mean, but writing it out helps you remember it is there.

Step 4: wrap the app and record some runs

app.py
bot = HRBot()

tru_bot = TruApp(
    bot,
    app_name="HR Bot",
    app_version="v1",
    feedbacks=[f_groundedness, f_answer_relevance, f_context_relevance],
)

questions = [
    "How many days of annual leave do I get?",
    "Can I carry unused leave into next year?",
    "What are the office hours?",
    "Do you offer a company car?",
]

for q in questions:
    with tru_bot as recording:
        bot.query(q)

print(session.get_leaderboard())

TruApp(...) wraps your object and registers it with the name HR Bot and the version v1. The argument that takes the metrics is still spelled feedbacks=, even though the objects are now Metrics; that is a leftover of the old name and it is not a bug in your code. The with tru_bot as recording: block is the recording context: everything your app does inside it is captured as one record. You call the original object, bot.query(q), not the wrapper, and TruLens observes from the side.

The last question is deliberate. The documents say nothing about company cars, so a good model should say it does not know. That is a useful test case, because it is exactly where groundedness matters: an answer of "Yes, all staff get a car" would be a hallucination, and the groundedness score should drop.

Run the file with python app.py. The metrics are computed in the background after each record is ingested, so the leaderboard printed at the very end may be empty or partial on a first run. That is not a failure; the next section explains why, and how to wait for the scores.

Scores arrive a moment after your code returns Evaluation runs on a background thread. If your script ends immediately after the last question, the process can exit before the last scores are written. Wait for results explicitly, as shown in the next section, instead of assuming they are ready.
Try it
  1. Build app.py from the four steps, using a key you control.
  2. Before adding TruLens, run HRBot().query(...) alone and confirm you get an answer.
  3. Add the instrumentation and metrics, run the file, and note any warnings printed.
an answer with no TruLens, then a run that prints the squid line and finishes without a traceback. If you see a missing-key error naming OPENAI_API_KEY, set the variable in the same terminal.

Waiting for scores and reading the leaderboard

Because evaluation is asynchronous, you need a way to say "hold on until the scores exist". TruLens gives you one on the recording object. Replace the final loop of your script with this:

app.py
for q in questions:
    with tru_bot as recording:
        bot.query(q)
    results = recording.retrieve_feedback_results(timeout=180)
    print(q)
    print(results.T)

session.force_flush()
print(session.get_leaderboard())

recording.retrieve_feedback_results(timeout=180) blocks until the metrics for that recording finish, up to 180 seconds, and returns a pandas DataFrame with one column per metric. session.force_flush() at the end pushes any spans still queued out to the database before the script exits. Use both habits in scripts; in a notebook you can usually rely on the cell finishing slowly enough, but a script cannot.

Do not use wait_for_feedback_results() on the app object. That method existed in older versions and in older tutorials, and under the current tracing it raises an error telling you to use retrieve_feedback_results instead.

Now read the output. Each score is a float between 0 and 1, where higher is better for every metric in the RAG Triad. A typical result for the first question might look like this:

TEXT
Groundedness        1.0
Answer Relevance    1.0
Context Relevance   0.67

How do you read that? Groundedness of 1.0 says every claim in the answer was supported by the retrieved text. Answer relevance of 1.0 says the answer addressed the question. Context relevance of about 0.67 says that, averaged over the two retrieved chunks, one was a strong match and the other was only partly relevant, which matches our toy retriever, which returns the top two keyword matches whether or not the second is any good. That is a real signal: it tells you retrieval, not generation, is the weakest link, and it points at the exact knob (how many chunks, or how they are ranked) to turn.

Where do the scores come from? The judge model is prompted to rate on a small integer scale, by default 0 to 3, and TruLens divides by the maximum to land in 0 to 1. So your scores cluster at values like 0, 0.33, 0.67 and 1, which are not measurements of a continuous quantity but the judge's coarse opinion. Averages over many records smooth this out; single scores on single records are noisy, and this is the place to start being sceptical.

session.get_leaderboard() aggregates across all the records of each app version and returns a DataFrame: one row per app version, with the mean of each metric, plus average latency and total cost. With only four questions it is a smoke test rather than a verdict, but its shape is the shape you will use all the time, for example to compare v1 and v2 of the bot on the same questions.

You can also fetch the raw records, which is how you look at individual runs from code:

PYTHON
records, feedback_columns = session.get_records_and_feedback(app_name="HR Bot")
print(feedback_columns)
print(records[["input", "output"] + feedback_columns].head())

get_records_and_feedback returns a DataFrame of records and a list naming the metric columns. Filtering by app_name and, if you like, app_version keeps it focused.

Try it
  1. Switch your loop to use retrieve_feedback_results and add session.force_flush() at the end.
  2. Run it and find the score for the company-car question.
  3. Change the system prompt to "Answer from your own knowledge" and run again under a new version label, v1-bad.
the company-car answer scores high on groundedness when the model says it does not know, and the deliberately bad prompt pulls groundedness down because the model now adds facts from outside the documents.

The dashboard: seeing runs and comparing versions

Numbers in a terminal are fine for a smoke test, but the point of TruLens is to explore. The dashboard is a local web page, built with Streamlit, that reads the same database your app wrote to. Start it from code, in the same folder:

PYTHON
from trulens.core import TruSession
from trulens.dashboard import run_dashboard

session = TruSession()
run_dashboard(session)

Or, since 2.14.0, from the shell, with no Python file at all:

BASH
trulens-dashboard

The command looks in the current directory for default.sqlite (or another SQLite file containing TruLens tables) and serves it. If it cannot find one, it says No TruLens database found and tells you to run your app first or pass --database-url. A few flags are worth knowing from day one: --port 8502 picks the port, --address 127.0.0.1 restricts it to your own machine, --force stops a dashboard that is already running, and --find prints which database it would serve and exits.

The dashboard prints a local URL, usually http://localhost:8501 or the port you chose. Open it. There are four pages.

Leaderboard lists your app versions with the mean of each metric, latency and cost. This is where you compare "HR Bot v1" with "HR Bot v1-bad" at a glance. Colours highlight the metrics so that weak spots stand out.

Records is where you learn the most. It lists every individual run with its input, output and scores. Click one and you get the trace: the tree of spans, showing the retrieval, the generation, and the timing of each, with the inputs and outputs of every step. Below it are the metric results with the judge's written reasons. When a score surprises you, this is where you find out why, by reading the reason the judge gave.

Compare puts two app versions side by side, record against record, so you can see how the answer to the same question changed.

Trends (added in 2.13) plots metrics, latency and cost over time, with confidence intervals. It becomes useful once you have a stream of production traffic rather than a handful of test questions.

The habit to build is this: never trust an aggregate number without opening a few of the records behind it. If the leaderboard says Groundedness is 0.55, open the lowest-scoring records, read the answer, read the retrieved context, read the judge's reason, and decide whether you agree. Sometimes the app is wrong. Sometimes the judge is. Learning to tell which is the whole craft of LLM evaluation, and the dashboard is where you practise it.

The dashboard has no login It is a plain Streamlit app with no authentication, and it shows prompts, documents and answers. On your laptop that is fine. Never expose it on a public address; bind it to 127.0.0.1 and put proper access control in front of it if a team needs to see it.

If the page says the dashboard is already running, you have a previous instance holding the port. Pass force=True (or --force on the command line) and it stops the old one first.

Try it
  1. Start the dashboard with trulens-dashboard from your project folder.
  2. On Records, open the company-car record and read the trace from the top.
  3. Find the judge's reason for the groundedness score and say whether you agree.
a span tree with a retrieval and a generation under the root, and a reason that cites the sentences it checked. Disagreeing with the judge at least once is a normal and healthy outcome.

Selectors and instrumentation in more depth

Most of your debugging in the first week will be about getting the right data into the right metric, so it is worth understanding the two halves of that pipeline properly: how data gets onto spans (instrumentation) and how metrics pull it off (selectors).

Putting data on spans

The @instrument decorator wraps a method. Every time the method runs inside a recording, TruLens opens a span, runs the method, and closes the span. Used with no arguments it still records something useful: each argument is stored as ai.observability.call.kwargs.<name> and the return value as ai.observability.call.return. The attributes argument lets you give those values standard names such as RETRIEVAL.QUERY_TEXT. The span_type argument tags the span so that selectors can find spans by kind.

There are three ways to decide what a span records:

  • a string naming a parameter, like "query", or the word "return";
  • a function that receives the return value, any exception, and the call arguments, and returns a dictionary of attributes, for when you need to transform a value first;
  • nothing, in which case you get the default capture above.

When the method you want to watch lives in a library you cannot edit, you do not decorate it; instead you patch it from outside with instrument_method(cls=..., method_name=..., span_type=..., attributes=...). You will not need it for your first app, but it is the answer when someone asks "how do I trace code I do not own?"

If your app uses LangChain, LangGraph or LlamaIndex, the framework wrappers (TruChain, TruGraph, TruLlama) instrument the framework's own steps for you, so you may not need any decorators at all. The concepts are identical; the wrapper just does the marking.

Pulling data off spans

A selector answers "which value on which span". The shortcuts cover nearly everything a beginner needs:

Selector Gets
Selector.select_record_input() the record root's input, usually the user's question
Selector.select_record_output() the record root's output, the final answer
Selector.select_context(collect_list=True) all retrieved contexts, as one list
Selector.select_context(collect_list=False) each retrieved context, scored one at a time

When a shortcut is not enough you build a selector from parts, for example by span type and attribute:

PYTHON
Selector(
    span_type=SpanAttributes.SpanType.RETRIEVAL,
    span_attribute=SpanAttributes.RETRIEVAL.QUERY_TEXT,
)

or by function name, when you want a value from one specific function of your own code:

PYTHON
Selector(function_name="app.HRBot.generate_completion", function_attribute="return")

Two failure patterns are worth memorising. If a metric never produces a score, first suspect that the span it needs does not exist or lacks the attribute: check that your retrieval method has the RETRIEVAL span type and the RETRIEVED_CONTEXTS attribute. If a metric raises an error about a missing argument, suspect a key in selectors that does not match a parameter name of the implementation. Between them, those two causes explain the large majority of "my metric does nothing" questions.

You can also write your own metric. The implementation can be any Python function that returns a number between 0 and 1. A classic first one is a length check:

PYTHON
f_short = Metric(
    name="Short Answer",
    implementation=lambda text: 1.0 if len(text.split()) <= 60 else 0.0,
    selectors={"text": Selector.select_record_output()},
)

It costs nothing and calls no model, which makes it a good way to experiment with selectors without spending tokens. Plain-code metrics like this are also the most trustworthy, because they are deterministic: the same answer always gets the same score.

Try it
  1. Add the Short Answer metric to the feedbacks list of your recorder under a new version, v2.
  2. Run your four questions again and open the leaderboard.
  3. Deliberately misspell the selector key ("txt") and run once more to see how the failure looks.
a fourth metric column appears with scores of 1.0 or 0.0, and the misspelled key produces an error or a missing column rather than a quiet success. Seeing the failure once makes it easy to recognise next time.

The provider methods you will reach for

The judge's built-in methods are grouped by what you want to measure. You do not need to memorise them, but knowing the families lets you choose without hunting through the reference. Each of these returns a float, and most have a _with_cot_reasons twin that returns the score plus an explanation.

Family Methods Use it to check
RAG Triad relevance, context_relevance, groundedness_measure_with_cot_reasons the three RAG failures
Quality coherence, conciseness, correctness, helpfulness, sentiment how the answer reads
Safety harmfulness, maliciousness, controversiality, misogyny, criminality, insensitivity, stereotypes unsafe or biased output
Citations citation_attribution, citation_accuracy whether cited sources back the claim
Conversation conversation_helpfulness, topic_adherence, coherence_across_turns multi-turn chats
Agents logical_consistency_with_cot_reasons, plan_adherence_with_cot_reasons, tool_selection_with_cot_reasons and similar an agent's whole trace

Two practical notes. First, the safety metrics score the presence of the bad property, so for some of them a high score is bad; check each method's docstring and, in any guardrail, mark such metrics as lower-is-better. Second, names change. In the past qs_relevance was renamed to relevance, and the agent metrics briefly carried a trajectory_ prefix that was later dropped, so an old tutorial's method name may simply not exist anymore. When a method is missing, look at the current provider reference rather than guessing.

The same provider object can drive any of them, so the pattern is always the same: pick a method, work out which parameters it takes, map a selector to each, give it a name, and add it to the recorder. Nothing else about your app changes.

Configuration and common errors

TruLens has very little configuration a beginner must touch, and most of what goes wrong is one of a dozen recurring errors. This section covers both, with the messages as they actually appear.

The settings you will actually use

Where data is stored. The default is the SQLite file default.sqlite. To use another file or database, pass a URL once, at the start of your program, before anything else creates the session:

PYTHON
from trulens.core import TruSession

session = TruSession(database_url="sqlite:///my_experiments.sqlite")

SQLite is ideal for learning and for a single developer, and unsuitable once several processes or machines write at the same time. The mid and senior guides cover PostgreSQL for that case.

Starting clean. session.reset_database() wipes all tables. It is convenient in tutorials and dangerous anywhere else, because it deletes your history without asking. Keep it out of any script you run against data you care about.

Redacting secrets. TruSession(database_redact_keys=True) masks API keys that might otherwise be captured inside stored objects. It is off by default; turn it on as soon as you share a database with anyone.

Turning tracing off. Setting the environment variable TRULENS_OTEL_TRACING=0 disables the OpenTelemetry tracing that everything in this guide depends on. You will rarely want to, and TruLens warns once if you do; it matters because a colleague's stray environment variable is a common reason for "no spans are recorded".

Errors, and how to read them

Key OPENAI_API_KEY needs to be set; please provide it in one of these ways... The judge has no credentials. Set the variable in the same terminal session you run Python from, or put it in a .env file. The message lists the options, so read it to the end.

These packages are required for using <X>: <package>. You should be able to install them with pip: pip install "<package>" You used an integration whose package is not installed, typically trulens-apps-langchain or a provider. Run the pip install line it prints.

Cannot use feedback_mode other than WITH_APP_THREAD with OTel tracing! You copied an older tutorial that passes feedback_mode=FeedbackMode.DEFERRED to the recorder. Delete the argument. Current TruLens computes metrics on a background thread attached to the app, and other modes are not available under the default tracing.

Already recording with a context manager, cannot nest! You opened a with tru_bot as recording: block inside another one. Use one recording context at a time and call your app once per block.

Cannot find TruLens context. The instrumented method ran on a different thread or asyncio task from the recording, so TruLens lost track of which record it belonged to. The fix is to run the work with TruLens's own thread helpers, or to keep it on the same thread. The doc page named in the message lists the options.

Database schema is behind the expected revision... You upgraded TruLens and pointed it at a database created by an older version. Back the file up, then run TruSession().migrate_database(). The mirror message, Database schema is ahead of the expected revision, means an older TruLens is reading a database a newer one already upgraded; the fix is to upgrade the older client.

TruSession was already initialized. Cannot change database configuration after initialization. You called TruSession(database_url=...) after a session already existed in the same process. The session is a singleton; configure it once, at the very top, and everything else reuses it.

DeprecationWarning: Feedback is deprecated... Use Metric instead Not an error, a signal. You, or a tutorial, used the old class. Replace Feedback(...) with Metric(implementation=..., selectors={...}) when convenient.

RuntimeError: Dashboard is already running An earlier dashboard still holds the port. Use force=True, or stop the old process.

Failures that print nothing

The most frustrating problems are the silent ones, and they share a short list of causes.

  • No scores, no errors. The selectors found nothing. Check the span type and attributes on retrieval, and the selector keys.
  • Scores missing for the last record. The script ended before the background thread finished. Use retrieve_feedback_results and force_flush.
  • An empty dashboard. You are in a different folder from the one that holds default.sqlite, or the dashboard is reading another file. Run trulens-dashboard --find to see which database it picked.
  • Scores look random. Your judge is too small or too noisy, or your questions are too few. Try a stronger judge and more examples before you change your app.
Try it
  1. Unset OPENAI_API_KEY in a fresh terminal and run your app, then read the full error.
  2. Set it again. Next, call TruSession(database_url="sqlite:///other.sqlite") after your first TruSession() and read the warning.
  3. Run trulens-dashboard --find from two different folders.
a missing-key error that lists how to fix it, a singleton warning that shows the second call was ignored, and two different answers to which database would be served. Each is a message you will meet again.

Two common mistakes about what the scores mean

Having the scores is not the same as understanding them, and two misunderstandings cause more bad decisions than any API slip.

The first is treating the judge as truth. Every LLM-judged metric is a language model's opinion, produced by a prompt. It can be inconsistent, it can be fooled by confident phrasing, and it has biases, such as favouring longer answers. A score of 0.67 is not a measurement of the answer's quality the way a thermometer reading is a measurement of temperature. It is one judge's coarse rating. The remedy is not to distrust everything but to calibrate: read a sample of the judge's reasons, compare them with your own judgement on a few dozen examples, and use scores mostly to compare versions, where consistent bias cancels out, rather than as absolute grades. The mid-level guide covers how to check a judge against human labels.

The second is drawing conclusions from tiny samples. Four questions can show you that the machinery works. They cannot tell you whether v2 is better than v1. A difference of 0.05 in an average over ten records is well inside the noise. Build an evaluation set of dozens to hundreds of realistic questions, including the nasty ones (questions the documents cannot answer, questions with typos, questions that try to get the bot to misbehave), and run every version against the same set. If you keep only one habit from this guide, keep that one: fixed question set, fixed judge, one change at a time.

A trustworthy comparison

  • The same 50 or more questions for v1 and v2
  • The same judge model, pinned by name
  • One change between versions
  • Lowest-scoring records read by a human

A misleading comparison

  • Four hand-picked questions
  • A judge left on its default and upgraded silently
  • Prompt, model and chunking all changed at once
  • Only the averages looked at

Putting it all together

Let us finish with one small project that uses everything: two versions of the HR bot, evaluated on the same fixed question set, compared in the dashboard. The change between versions is one line, the system prompt, so any difference in scores has one explanation.

compare.py
import numpy as np
from openai import OpenAI as OpenAIClient
from trulens.apps.app import TruApp
from trulens.core import Metric, Selector, TruSession
from trulens.core.otel.instrument import instrument
from trulens.otel.semconv.trace import SpanAttributes
from trulens.providers.openai import OpenAI

DOCS = [
    "Employees receive 25 days of paid annual leave per calendar year.",
    "Unused leave can be carried over, up to a maximum of 5 days.",
    "The office is open Sunday to Thursday, from 9 am to 5 pm.",
    "Remote work is allowed up to two days per week with manager approval.",
]
QUESTIONS = [
    "How many days of annual leave do I get?",
    "Can I carry unused leave into next year?",
    "What are the office hours?",
    "How many days can I work remotely?",
    "Do you offer a company car?",
    "What is the dental insurance limit?",
]

client = OpenAIClient()


class HRBot:
    def __init__(self, system_prompt: str):
        self.system_prompt = system_prompt

    @instrument(
        span_type=SpanAttributes.SpanType.RETRIEVAL,
        attributes={
            SpanAttributes.RETRIEVAL.QUERY_TEXT: "query",
            SpanAttributes.RETRIEVAL.RETRIEVED_CONTEXTS: "return",
        },
    )
    def retrieve(self, query: str) -> list[str]:
        words = {w.lower().strip("?.,") for w in query.split()}
        scored = [(len(words & set(d.lower().replace(".", "").split())), d) for d in DOCS]
        scored.sort(reverse=True)
        return [d for score, d in scored[:2] if score > 0]

    @instrument(span_type=SpanAttributes.SpanType.GENERATION)
    def generate_completion(self, query: str, context_list: list[str]) -> str:
        context = "\n".join(context_list)
        response = client.chat.completions.create(
            model="gpt-4o-mini",
            messages=[
                {"role": "system", "content": self.system_prompt},
                {"role": "user", "content": f"Context:\n{context}\n\nQuestion: {query}"},
            ],
        )
        return response.choices[0].message.content

    @instrument(
        span_type=SpanAttributes.SpanType.RECORD_ROOT,
        attributes={
            SpanAttributes.RECORD_ROOT.INPUT: "query",
            SpanAttributes.RECORD_ROOT.OUTPUT: "return",
        },
    )
    def query(self, query: str) -> str:
        return self.generate_completion(query, self.retrieve(query))


session = TruSession()
provider = OpenAI(model_engine="gpt-4o-mini")

metrics = [
    Metric(
        implementation=provider.groundedness_measure_with_cot_reasons_consider_answerability,
        name="Groundedness",
        selectors={
            "source": Selector.select_context(collect_list=True),
            "statement": Selector.select_record_output(),
            "question": Selector.select_record_input(),
        },
    ),
    Metric(
        implementation=provider.relevance_with_cot_reasons,
        name="Answer Relevance",
        selectors={
            "prompt": Selector.select_record_input(),
            "response": Selector.select_record_output(),
        },
    ),
    Metric(
        implementation=provider.context_relevance_with_cot_reasons,
        name="Context Relevance",
        selectors={
            "question": Selector.select_record_input(),
            "context": Selector.select_context(collect_list=False),
        },
        agg=np.mean,
    ),
]

VERSIONS = {
    "strict": "Answer using only the context. If the answer is not in it, say you do not know.",
    "loose": "You are a helpful HR assistant. Answer the question as best you can.",
}

for version, prompt in VERSIONS.items():
    bot = HRBot(prompt)
    tru_bot = TruApp(bot, app_name="HR Bot", app_version=version, feedbacks=metrics)
    for q in QUESTIONS:
        with tru_bot as recording:
            bot.query(q)
        recording.retrieve_feedback_results(timeout=180)

session.force_flush()
print(session.get_leaderboard())

Run it with python compare.py, then start trulens-dashboard in the same folder. On the leaderboard you should see two rows, strict and loose. The instructive rows are the two questions the documents cannot answer, the company car and the dental limit. The strict prompt should say it does not know, which keeps groundedness high. The loose prompt invites the model to improvise, and any invented figure will lower its groundedness. Open the Records page for those two questions under each version, read the answers side by side on the Compare page, and read the judge's reasons. You have now done, in miniature, the complete loop: instrument, record, evaluate, compare, iterate.

A note on cost and honesty. Each metric per record is at least one judge call, and context relevance with collect_list=False is one call per chunk. Six questions, three metrics, two versions is a few dozen calls, which costs cents. The same script over a thousand questions is a real bill, and later guides show how to sample. Also, your exact numbers will differ from any written here, because models vary from run to run. Judge the pattern, not the digits.

Try it
  1. Run compare.py and open the dashboard.
  2. On the leaderboard, say which version wins on each metric.
  3. Open the dental-insurance record under both versions and decide whether the judge's groundedness reasoning matches your own.
two versions that differ mostly on groundedness for the unanswerable questions, and at least one place where you can explain in words why a score is what it is. That explanation is the skill; the number is only the prompt for it.

What you can now do, and what comes next

You can now explain what TruLens is for and how its five nouns fit together: session, app, record and span, metric, provider. You can install the current packages, pin a judge, instrument a small app with @instrument, define the three RAG Triad metrics with selectors, record runs, wait for and read the scores, and open the dashboard to compare two versions and inspect a single trace. You can recognise the usual errors by their messages, you know the silent failures and where to look for them, and you know not to treat a judge's score as truth or a four-question sample as proof.

What comes next is depth. The Mid-level guide shows how the tracing and the background evaluator actually work, how to move storage to PostgreSQL, how to evaluate a whole dataset offline with BatchEvaluator and Runs, how to write guardrails that block bad inputs and outputs at runtime, how to evaluate conversations and agents, and how to check a judge against human labels. The Senior guide treats TruLens as a platform: sampling and cost budgets for online evaluation, exporting to an OpenTelemetry backend, security of the dashboard and of captured content, upgrades and database migrations, multi-team isolation, and when to choose a different tool.

If you want to look sideways first, read the guides for Langfuse (/student-guides/langfuse) and RAGAS (/student-guides/ragas) to see how other tools approach tracing and scoring, and the guide for LangGraph (/student-guides/langgraph) if your app is an agent. Comparing the same idea in two tools is the fastest way to learn which parts are essential and which are one tool's accent.

Sources