Skip to content
Back to student guides
Arize PhoenixLLMOpsObservability & tracing3 levels106 sectionsCovers Arize Phoenix 20.16

The Complete Arize Phoenix Guide

Trace and evaluate LLM apps with the open-source Arize Phoenix, built on OpenTelemetry. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
17sections
27examples

This is part one of three. It covers everything you need to do real work with Arize Phoenix, not a teaser. By the end you can start a Phoenix server on your own laptop, send traces from a Python LLM application into it, read those traces in the web interface, label them by hand, score them automatically with a second model acting as a judge, save interesting cases into a dataset, and run an experiment against that dataset. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas stick only once you have watched your own application appear, span by span, in a browser tab. The examples use Python and the OpenAI client because they are the most common starting point, but the ideas carry over to every provider and framework Phoenix supports.

One honest warning first. Phoenix ships a new minor release every few days, and this guide was checked against version 20.16.0, released on 23 September 2026. The version you install tomorrow will be newer. Most of what you read here is stable because it rests on OpenTelemetry, but the Phoenix-specific APIs did change a lot between versions 12 and 14, and many tutorials on the internet still show the old forms. Where that matters, this guide says so and shows the current form.

What Phoenix is, and the problem it solves

Arize Phoenix is an open-source platform for seeing inside applications that use large language models, and for measuring whether they behave well. You run it as a small server. Your application sends it a record of everything it did on each request: which prompt went to which model, what came back, how many tokens it cost, how long each step took, and which tool or database lookup happened in between. Phoenix stores those records, draws them as a tree in a web interface, and gives you tools to label them, score them, and turn them into test cases.

YOUR APPcalls an LLM
→
TRACESa record per request
→
PHOENIXstores and shows them
→
EVALUATElabel, score, test

To see why this is needed, think about what a normal web service gives you when something goes wrong. You have logs, a stack trace, and a status code. An LLM application is different in a way that breaks all three. The call succeeds, the status code is 200, no exception is raised, and the answer is still wrong. The model invented a policy that does not exist, ignored the document you retrieved, or called the wrong tool with plausible-looking arguments. Nothing in a conventional log line tells you that. To find out, you have to read the actual prompt and the actual reply, and you have to read them in the context of the steps around them.

Applications also stopped being a single call. A modern assistant may rewrite the user's question, search a vector database, pass the top five passages to a model, let the model call a calculator tool, feed the result back, and finally produce an answer. That is six or seven operations, several of them non-deterministic, and a bad answer could be caused by any of them. If the retrieval step returned the wrong passages, rewriting the final prompt will not help. Without a record of every step you are guessing.

Phoenix solves this in two halves that feed each other.

Observability is the first half: a trace of each request, captured automatically, so you can open any single interaction and see exactly what happened. This is how you debug.

Evaluation is the second half: a way to attach judgements to those interactions, whether a person clicked "thumbs down", a rule checked a format, or another model scored the answer for faithfulness. Judgements let you move from "this one looks wrong" to "eleven percent of answers are unfaithful, and the rate went up after Tuesday's change". This is how you improve.

The loop the Phoenix documentation teaches is the same one this guide follows: trace, then annotate or evaluate, then curate datasets, then run experiments, then iterate on your prompts, then trace again. Each arrow in that loop is a section below.

What came before

Teams first reached for print statements and log files. Prompts and completions were written to a log line, and someone searched with grep. That works for ten requests and fails for ten thousand, because a flat log has no structure: you cannot ask it to show the slowest step, or all requests where the tool call failed, or the tree of operations behind one particular answer.

Next came general application performance monitoring tools. They do understand requests as trees of timed operations, but they were built for latency and errors, not for the content of a prompt. They show that a call took 2.1 seconds. They do not show the text, do not count tokens by model, and have no concept of scoring an answer.

Then came a wave of LLM-specific platforms, a category usually called LLM observability or LLMOps tooling. Phoenix belongs there, next to Langfuse, LangSmith and others. Its particular choice is to be built directly on OpenTelemetry, the open standard that the rest of the observability world uses for traces. That choice has a practical consequence: the data your application emits is not locked to Phoenix. You can send the same traces to Phoenix and to another backend at once, and you can keep your instrumentation if you ever switch.

Phoenix also used to be a tool for classic machine-learning model monitoring, with embedding visualisations and drift analysis. That functionality was removed in version 13. If an older course or blog post shows px.Inferences, embedding point clouds or a Schema object, it is teaching a version of Phoenix that no longer exists. This guide teaches only the LLM and agent side, which is what Phoenix is today.

Three products with similar names

Newcomers often confuse the offerings, so separate them now.

  • Phoenix is the open-source project this guide is about. You install it yourself and run it wherever you like. It is licensed under the Elastic License 2.0, which is free to use and to self-host but is not an Apache or MIT licence, so read it if your employer has a policy on licences.
  • Phoenix Cloud is the same Phoenix hosted for you at app.phoenix.arize.com. You sign up, receive an endpoint that looks like https://app.phoenix.arize.com/s/<your-space>, and skip the server setup.
  • Arize AX is Arize's commercial enterprise platform, with multi-team organisations and managed hosting. It is a separate product.

For a student or a first project, running Phoenix on your own laptop is the right choice: it costs nothing and needs no account. For employers in the Gulf or Egypt with data-residency rules, the self-hosted route has a real advantage too. Prompts and completions often contain customer data, and when you run Phoenix yourself, that data stays in the database you chose, in the cloud region you chose, rather than leaving for a hosted service.

Try it
  1. Pick any LLM feature you have used or built, such as a chatbot or a document summariser.
  2. Write down the last time it gave a wrong answer. List every step that happened between the question and the answer.
  3. For each step, write what you would need to see to know whether that step was the culprit.
a list of things like "the passages retrieved", "the exact prompt after templating" and "which tool was called". Every one of those is something a trace records automatically, and the reason to learn this tool.

The mental model: four nouns

Phoenix has a large surface, but the core vocabulary is small. Learn these four nouns well and every screen will make sense.

Trace and span

A span is one unit of work. Calling the model is a span. Running a search is a span. A Python function you decorated is a span. Each span records a name, a start time and an end time, a status (OK, ERROR or UNSET), a pointer to its parent span, and a bag of key-value attributes describing what it did.

A trace is the complete record of one request moving through your application. Technically it is the set of all spans that share the same trace_id, arranged as a tree. The first span has no parent and is called the root span. Everything it triggered hangs beneath it.

ROOT SPANanswer_question
→
CHILDretrieve_documents
→
CHILDChatCompletion (LLM)

If you have used a browser's network tab or a profiler's flame graph, you already know the shape. The tree tells you what called what, and the time bars tell you where the seconds went.

Attributes and span kind

Attributes are where the LLM-specific information lives. For a model call, Phoenix expects attributes such as llm.model_name, input.value, output.value, and the token counts llm.token_count.prompt and its completion counterpart. You rarely write these by hand. The instrumentation library fills them in.

The set of names comes from OpenInference, a set of conventions that Arize maintains on top of OpenTelemetry. One of those conventions is the span kind, stored in the attribute openinference.span.kind. It tells Phoenix what sort of work the span represents, so the interface can show a retriever's documents differently from a model's messages. The kinds are written in capitals: LLM, CHAIN, AGENT, TOOL, RETRIEVER, EMBEDDING, RERANKER, GUARDRAIL, EVALUATOR, PROMPT and UNKNOWN. A span with no kind set is UNKNOWN.

Project

A project is a container of traces, normally one per application or environment. When you open Phoenix you first see a list of projects. The default project is literally named default, and you should avoid using it for anything real, because traces from every unnamed experiment pile up there. You choose the project from your code with project_name, which Phoenix writes onto your data as a resource attribute called openinference.project.name.

Instrumentation

Instrumentation is the code that creates spans. There are two styles. Automatic instrumentation uses a library that knows how to wrap a framework: install the OpenAI instrumentor and every OpenAI call becomes an LLM span without changing your application code. Manual instrumentation is where you mark your own functions, using decorators that Phoenix provides or the standard OpenTelemetry tracer. Real applications use both: automatic for the model and framework calls, manual for your own business logic.

Sessions, annotations, datasets, experiments

Four more nouns appear later in the guide, and it helps to meet them briefly now.

  • A session groups several traces that belong to one conversation, using a shared session.id. The interface shows it as a chat timeline.
  • An annotation is a label, a score, or a short explanation attached to a span or a trace. Its source, the annotator_kind, is HUMAN, LLM or CODE. You will see older material call these "evaluations"; since version 14 the word is annotations everywhere.
  • A dataset is a versioned collection of examples, each with an input and usually an expected output.
  • An experiment runs a function, called the task, over every example in a dataset and scores the results with evaluators.
OpenTelemetry in one sentence OpenTelemetry (OTel) is the open standard for producing and shipping traces, metrics and logs. Phoenix does not invent a proprietary format: it is an OTel-compatible receiver that understands the extra LLM attributes from OpenInference. That is why the trace code you write here is portable.
Try it
  1. Sketch, on paper, the trace for this request: "What is our refund policy?" to a bot that searches documents, then asks a model.
  2. Draw one root span with two children. Label each with a span kind from the list above.
  3. Add two attributes to each child that you would want to see.
a root of kind CHAIN or AGENT, a RETRIEVER child with the retrieved documents, and an LLM child with the prompt and the reply. You have just designed what Phoenix will display.

Installing the server and checking it works

Phoenix is an ordinary Python package, so you install it the way you install any Python tool. It supports Python 3.10 through 3.14 and runs on Linux, macOS and Windows. Always use a virtual environment, a private folder of packages for one project, so that Phoenix's many dependencies cannot collide with anything else on your machine.

On Linux or macOS:

BASH
python3 -m venv .venv && source .venv/bin/activate
pip install arize-phoenix
phoenix serve

On Windows, in PowerShell:

POWERSHELL
py -m venv .venv
.\.venv\Scripts\Activate.ps1
pip install arize-phoenix
phoenix serve

The first two commands create and enter the virtual environment, pip install downloads the server and its web interface, and phoenix serve starts it. The terminal will print a startup log and stay busy. That is correct: a server is a program that keeps running until you stop it with Ctrl+C. Leave this terminal open and use a second one for everything else.

If you would rather install nothing permanently, uvx arize-phoenix serve runs the server through uv, a fast Python package runner, without a separate install step. The documentation calls it the quickest local option.

Notice the order of the words. Flags go after serve. Before version 14 the command line was arranged differently, and old tutorials show phoenix --port 6006 serve, which now fails. The current form is:

BASH
phoenix serve --host 0.0.0.0 --port 6006

What the server listens on

Phoenix opens two ports that matter now.

Port What it carries
6006 The web interface, the REST and GraphQL APIs, and trace ingest over HTTP at /v1/traces
4317 Trace ingest over gRPC, a second way for your application to send spans

Why two? OpenTelemetry defines two transports for sending traces, HTTP with protobuf encoding and gRPC, and Phoenix accepts both. Which one your application uses will matter in a later section, because mixing them up is the most common first-day failure.

Where your data lives

With no database configured, Phoenix stores everything in a SQLite file inside a working directory, which defaults to .phoenix in your home folder. That means your traces survive restarting the server, and you can delete the folder to start completely fresh. SQLite is perfect for learning and for development. A shared or production Phoenix should use PostgreSQL instead; Mid-level and Senior cover that.

If you work in a Jupyter notebook you may also see import phoenix as px; px.launch_app(). That starts a server inside the notebook process, and it does not keep data by default. It is handy for a demo and a poor habit for real work. Prefer the separate phoenix serve process.

Checking that it is alive

Open http://localhost:6006 in a browser. You should see the Phoenix interface with an empty list of projects. From a second terminal you can also ask the server directly:

BASH
curl http://localhost:6006/healthz
curl http://localhost:6006/readyz
curl http://localhost:6006/arize_phoenix_version
pip show arize-phoenix

/healthz answers whether the process is alive. /readyz answers whether it is ready to do work, which includes being able to reach its database. The third returns the running server's version, and pip show reports the version of the package installed in your environment. If those two numbers differ, you are talking to a different server than the one you think you installed, a mix-up that costs people real time later.

Running it in Docker instead

If you already use Docker, the official image runs the same server:

BASH
docker run -p 6006:6006 -p 4317:4317 -i -t arizephoenix/phoenix:20.16.0

The -p flags publish the two ports so your host can reach them. Pin the version number rather than using latest; the project's documentation warns against latest for anything you care about. Data written inside this container disappears when the container is removed, unless you mount a volume and set PHOENIX_WORKING_DIR to point at it. If you do not know Docker yet, our Docker guide is the natural companion, and our Kubernetes guide covers the eventual production home.

Port already in use

If the server exits at startup with "Address already in use" (the error number differs by operating system: [Errno 48] on macOS, [Errno 98] on Linux, [WinError 10048] on Windows), something else is holding port 6006 or 4317, often an earlier copy of Phoenix you forgot about. Either stop that process or move Phoenix:

BASH
phoenix serve --port 6007 --grpc-port 4318

Remember the new ports when you point your application at the server.

A fresh install has no authentication The bare server starts with login turned off, which is fine on your laptop and dangerous anywhere reachable from the internet. Anyone who can reach the port can read every prompt and completion. Keep learning instances on localhost, and read the configuration section before you expose one.
Try it
  1. Create a virtual environment, install arize-phoenix, and run phoenix serve.
  2. Open http://localhost:6006 and confirm the interface loads.
  3. In a second terminal, run the three curl commands and compare the version with pip show arize-phoenix.
an empty project list in the browser, a healthy response from /healthz and /readyz, and two matching version numbers. You now have a working observability server.

Your first traced application

Now the main event. We will write a tiny program that asks an OpenAI model a question, then watch it appear in Phoenix. You need an OpenAI API key for this exact example, exported as OPENAI_API_KEY. Any other provider works the same way with its own instrumentor; we use OpenAI because it is the most widely documented.

Install the tracing helper and the OpenAI instrumentor into the same virtual environment as your application:

BASH
pip install arize-phoenix-otel openinference-instrumentation-openai openai

arize-phoenix-otel is a thin wrapper that configures OpenTelemetry for Phoenix with sensible defaults. openinference-instrumentation-openai is the piece that knows how to wrap the OpenAI client and turn each call into an LLM span. One version note: from release 0.17.2, arize-phoenix-otel needs Python 3.11 or newer, and on Python 3.10 pip quietly installs the previous release instead, which still works.

Create a file named first_trace.py:

first_trace.py
from phoenix.otel import register
from openai import OpenAI

tracer_provider = register(
    project_name="my-first-app",
    auto_instrument=True,
    protocol="http/protobuf",
    endpoint="http://localhost:6006/v1/traces",
)

client = OpenAI()

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Explain what a span is in one sentence."}],
)
print(response.choices[0].message.content)

Run it with python first_trace.py. Before the answer prints, register() writes a short banner describing what it configured: the project name, the span processor, the collector endpoint and the transport. Read that banner. It is the quickest diagnostic you will ever get, and later sections rely on it.

What each line did

register() is the one function you call at startup. It builds an OpenTelemetry TracerProvider, the object that creates spans and hands them to an exporter, and points it at Phoenix. Its arguments are keyword-only. The ones you will use first are:

Argument Meaning
project_name Which Phoenix project receives the traces. If omitted it reads PHOENIX_PROJECT or PHOENIX_PROJECT_NAME, and finally falls back to default.
auto_instrument When true, instruments every installed openinference-instrumentation-* package automatically.
endpoint Where to send spans.
protocol "http/protobuf" or "grpc".
batch When true, spans are buffered and sent in groups. Recommended for production. The default is false.
api_key For servers with authentication turned on. Also read from PHOENIX_API_KEY.

With auto_instrument=True, register() looks at what is installed and instruments it. If you prefer to be explicit, you can leave it off and call the instrumentor yourself:

PYTHON
from openinference.instrumentation.openai import OpenAIInstrumentor

OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)

Both routes give the same result. Explicit is clearer when you read the code a year later; automatic is shorter for a demo.

Why we passed the protocol and endpoint

We spelled out protocol="http/protobuf" and a full /v1/traces URL on purpose. If you pass a bare base URL such as http://localhost:6006 and say nothing about the protocol, the Python SDK infers gRPC on port 4317, not HTTP on 6006. That works when you started Phoenix locally and both ports are open. It fails silently when you later point at a server behind a load balancer that exposes only HTTPS. Being explicit from day one avoids learning this the hard way. The errors section returns to it.

Seeing the result

Switch to the browser and refresh. A project named my-first-app should now appear. Open it, and you will see a single row in the traces table. Click the row.

Try it
  1. Install the three packages, export OPENAI_API_KEY, and create first_trace.py.
  2. Run it and read the register() banner. Find the project name and the endpoint in it.
  3. Refresh Phoenix, open my-first-app, and click the trace.
a trace with one LLM span showing your prompt, the model's reply, the model name and the token counts. If the project never appears, jump to the errors section and start with the banner.

Reading a trace in the interface

A trace you cannot read is decoration. Spend this section learning the screen, because you will spend most of your working life in it.

Open your project. The default view is a table of traces, newest first, with columns for the start time, the input and output of the root span, the latency, the token counts and often a cost figure. Phoenix calculates cost per span from a built-in table of model prices, which it updates almost every release, so the numbers are estimates and you should check them against your provider's invoice before quoting them.

Click any row and a panel opens with the tree on the left and the details of the selected span on the right. Learn to read four areas.

The tree. Each indented row is a span, labelled with its name and a small icon for its kind. A long bar beside it shows how much of the total time it consumed. Bars that start late show work waiting on something earlier. In a real application, the tree is where you notice that retrieval took three seconds and the model call took one, a fact no amount of staring at the final answer would reveal.

The Info tab. For an LLM span this shows the input messages and output messages as a readable chat, including the system prompt, exactly as the model saw them after templating. This is the single most valuable view. Most "why did it say that" questions are answered by reading the assembled prompt and discovering it contained something you did not intend.

The Attributes tab. The raw key-value list behind the nice display. When something is missing from the pretty view, it is usually here in raw form. It also teaches you the attribute names, which you need when you write filters.

The status and events. A span that raised an exception has status ERROR, and the exception message and stack trace are recorded as an event on the span. A failing tool call in an agent is found by looking for the red span, not by reading a log.

Filtering the table

As traces multiply you will not scroll. The search bar above the table accepts a filter expression, a small Python-like language over span fields. A few you will use immediately:

TEXT
span_kind == 'LLM'
'refund' in input.value
status_code == 'ERROR'
parent_id is None

The first keeps only model calls, the second finds spans whose input contains a word, the third shows failures, and the last keeps only root spans, which gives you one row per request. Later sections add filters on annotations and metadata. The same expressions work in code, so the skill transfers.

Always start from the prompt When an answer is bad, open the LLM span and read the input messages before anything else. Nine times out of ten the cause is visible there: an empty retrieval result pasted into the prompt, a truncated document, or an instruction that contradicts another.
Try it
  1. Change the question in first_trace.py three times and run it each time.
  2. In Phoenix, filter with span_kind == 'LLM' and then with 'span' in input.value.
  3. Open one trace and find the prompt tokens, completion tokens and latency.
three rows, filters that narrow them as expected, and numbers that change when you ask for longer answers. Tokens and time are the two costs you will watch for the rest of your career with these systems.

Tracing your own code: decorators and span kinds

Automatic instrumentation captures calls to libraries it recognises. It knows nothing about your function look_up_policy() or the little routine that decides which department a ticket belongs to. Yet those are often exactly the steps you need to see. Manual instrumentation fills the gap.

The tracer you get from register() provides decorators. You create a tracer once, then mark functions:

support_bot.py
from phoenix.otel import register
from openai import OpenAI

tracer_provider = register(
    project_name="support-bot",
    auto_instrument=True,
    protocol="http/protobuf",
    endpoint="http://localhost:6006/v1/traces",
)
tracer = tracer_provider.get_tracer(__name__)
client = OpenAI()

POLICIES = {
    "refund": "Refunds are available within 14 days of purchase with a receipt.",
    "shipping": "Standard shipping to the Gulf region takes 5 to 8 working days.",
}


@tracer.tool
def look_up_policy(topic: str) -> str:
    """Return the policy text for a topic, or an empty string."""
    return POLICIES.get(topic, "")


@tracer.chain
def answer_question(question: str) -> str:
    topic = "refund" if "refund" in question.lower() else "shipping"
    policy = look_up_policy(topic)
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": f"Answer using only this policy: {policy}"},
            {"role": "user", "content": question},
        ],
    )
    return response.choices[0].message.content


if __name__ == "__main__":
    print(answer_question("Can I get a refund after two weeks?"))

Run it. In Phoenix you now see a single trace with a tree: a CHAIN span named answer_question at the root, a TOOL span for look_up_policy beneath it, and the automatic LLM span for the OpenAI call beside it. The decorator records the function's arguments as the span's input and its return value as the output, which is why the trace table shows a readable question and answer on the root row.

The decorators mirror the span kinds: @tracer.chain for a pipeline step, @tracer.tool for something an agent would call, @tracer.agent for an autonomous loop and @tracer.llm for a model call you wrap yourself. Pick the kind that matches what the function does, because the interface and the evaluators use the kind to decide how to display and treat the span. A retrieval function that returns documents is a RETRIEVER, and Phoenix will show the returned documents as a list with their scores, which is far more useful than a raw dump.

If you need full control, the standard OpenTelemetry tracer API also works, and you may use a with tracer.start_as_current_span("name") as span: block, then set attributes on the span yourself. Decorators are enough for the vast majority of beginner work.

Why the tree matters more than the flat list

Consider what you now gain. If the answer says "refunds take 30 days", you open the trace and see in the TOOL span that the policy text returned was empty, because the topic guess was wrong. The model did its best with nothing. The bug is in your keyword routing and not in the prompt, and you found that in under a minute. A flat log of only the model call would have sent you off rewriting the prompt for an afternoon.

Pausing and stopping tracing

Sometimes you do not want spans. A health check that calls a model every ten seconds would fill your project with noise. Wrap the code in a context manager to suppress tracing inside it:

PYTHON
from phoenix.trace import suppress_tracing

with suppress_tracing():
    client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "ping"}],
    )

To stop an instrumentor for good, call its uninstrument() method, for example OpenAIInstrumentor().uninstrument().

Streaming and token counts

A detail that catches people: when you stream a response from OpenAI, the provider does not include token usage in the stream unless you ask for it. The trace then shows no token counts and no cost. Pass stream_options={"include_usage": True} to the call and the numbers return.

Do not trace secrets Everything in a span's input and output is stored in your Phoenix database in readable form. A decorated function whose argument is a password, a national ID number or a customer's phone number will record it. Decorate functions whose inputs are safe to store, or keep sensitive values out of the arguments.
Try it
  1. Copy support_bot.py and run it with a refund question, then a shipping question, then a question about neither.
  2. For each trace, open the TOOL span and read what the lookup returned.
  3. Change the tool's decorator to @tracer.chain and compare how the span is labelled.
a three-level tree, a tool span whose output is empty for the third question, and a different icon after the decorator change. The empty lookup explains the weak answer without touching the prompt.

Sessions, users, metadata and tags

One trace is one request. Real products have conversations, users and versions, and you will want to slice your data by them. OpenInference provides context managers, wrappers that attach extra attributes to every span created inside them, so you write the identifier once at the boundary instead of passing it through every function.

chat_session.py
from openinference.instrumentation import using_attributes

def handle_message(conversation_id: str, user_id: str, text: str) -> str:
    with using_attributes(
        session_id=conversation_id,
        user_id=user_id,
        metadata={"app_version": "1.4.0", "channel": "web"},
        tags=["beta"],
    ):
        return answer_question(text)

Inside that block every span, including those created automatically by the OpenAI instrumentor, carries session.id, user.id, metadata and tag.tags. You can also use the smaller managers individually: using_session(session_id=...), using_user(user_id=...), using_metadata({...}) and using_tags([...]).

What sessions give you

Open your project and switch to the Sessions view. Phoenix groups traces that share a session.id and shows them as a chat timeline, turn by turn. That is how you judge a conversation rather than an isolated answer. A reply that looks fine alone can be plainly wrong in context, for instance when the bot forgets what the user said two turns earlier, and only the session view shows it.

Metadata and tags for slicing

Metadata is a dictionary of anything you like, and tags are a list of short labels. Their real value comes later, when you filter:

TEXT
metadata['channel'] == 'web'
metadata['app_version'] == '1.4.0'

When you release version 1.5.0 and complaints rise, you can compare the two versions directly. If you never recorded the version, you can only guess. The habit to adopt on day one is to stamp every trace with the version of your code and of your prompt.

A small caution on naming: be consistent. app_version in one place and appVersion in another produce two separate fields, and your filter will silently match only half the data.

Stamp the version early The moment you have a second version of your prompt, you need to know which one produced which trace. Add prompt_version to metadata from the start. It costs one line and saves the question "did this get worse after we changed it?" from being unanswerable.
Try it
  1. Wrap answer_question calls in using_attributes with the same session_id for three calls, then a different one for two more.
  2. Open the Sessions view and open each session.
  3. Filter the trace table with metadata['channel'] == 'web'.
two sessions, three turns and two turns, displayed as conversations, and a filter that returns exactly the traces you tagged.

Annotations: attaching judgements to traces

Seeing a trace tells you what happened. An annotation records what you think about it. An annotation is a named judgement attached to a span, a trace, a document retrieved inside a span, or a whole session. It can hold a label such as correct or incorrect, a numeric score, a free-text explanation, or any combination. It also records who made the judgement, in the annotator_kind: HUMAN for a person, LLM for a model acting as judge, or CODE for a rule.

The vocabulary changed in version 14. Older tutorials call annotations "evaluations" and use px.log_evaluations and SpanEvaluations. Those are removed. Everything is now annotation, and the old REST endpoints under /v1/evaluations no longer exist; the current ones are /v1/span_annotations, /v1/trace_annotations and /v1/document_annotations.

Annotating by hand in the interface

Open a trace, select a span, and use the annotation panel to add a label, a score or a note. A person reviewing twenty traces for ten minutes and marking each good or bad with a short reason produces surprisingly valuable data. It tells you what "good" means for your product, in a form you can later teach to an automatic judge.

To make annotation consistent, an admin can define an annotation config: a named, typed annotation that appears as a ready-made control. It is either categorical (a fixed set of labels), continuous (a number within a range) or freeform. Using a config means two reviewers cannot invent correct and Correct and right for the same idea.

There is also a special annotation called a note, free text attached to a span or a trace. The name note is reserved for these. Phoenix rejects an ordinary annotation named note with an error that tells you to use the notes endpoint.

Annotating from code

When feedback arrives from your own product, such as a thumbs-down button, you send it with the Python client. Install it and create a client:

BASH
pip install arize-phoenix-client
annotate.py
from phoenix.client import Client

client = Client(base_url="http://localhost:6006")

client.spans.add_span_annotation(
    span_id="PASTE_A_SPAN_ID_HERE",
    annotation_name="user_feedback",
    annotator_kind="HUMAN",
    label="thumbs_down",
    score=0.0,
    explanation="The bot quoted the wrong refund window.",
)

Two points deserve attention. First, import from phoenix.client. The old px.Client() and phoenix.session.client.Client were removed in version 14 and raise ModuleNotFoundError. Second, the parameter is base_url, which was called endpoint in the old client. The span id comes from the span panel in the interface, or from your own code if you capture it when the span is created.

Why annotations are the hinge of the whole loop

Everything later depends on them. Filter for annotations['user_feedback'].label == 'thumbs_down' and you have a worklist of problem cases. Put those cases into a dataset and you have a regression test. Ask an LLM judge to reproduce your human labels and you can measure the judge before you trust it. Annotation is not paperwork; it is the currency the rest of Phoenix spends.

Try it
  1. In the interface, annotate five of your traces by hand, each with a label and a one-line explanation.
  2. Filter the table with annotations['your_annotation_name'].label == 'bad', using the name you chose.
  3. Copy a span id and add one annotation with the Python client.
a filtered list containing exactly the traces you marked bad, and a sixth annotation on a span that you created from code. Humans and programs write to the same place.

Evaluations: letting a model score the answers

Reading traces by hand does not scale past a few dozen. To score hundreds or thousands you need an automatic judge. Phoenix's evaluation library, arize-phoenix-evals, provides evaluators: objects that take some input and return a score with a name, a label, a numeric value and an explanation. There are two families.

Code evaluators are plain functions with deterministic rules. Did the answer contain valid JSON? Did it match the expected string exactly? Is it shorter than 200 words? They are cheap, fast and reliable, and you should reach for them first wherever a rule can express the check.

LLM-as-a-judge evaluators ask a second model to grade the first. You write a prompt describing the criterion, for example "Is this answer fully supported by the provided context?", and constrain the judge to a short list of labels. They handle fuzzy qualities like faithfulness, relevance and tone, at the cost of money, latency and the judge's own mistakes.

Installing and a first judge

BASH
pip install arize-phoenix-evals openai pandas

The current evals API, introduced in version 2 and cleaned up in version 3.0 of the package, is built around an LLM object and a ClassificationEvaluator. The old names llm_classify, OpenAIModel and run_evals are gone. If you meet them in a tutorial, the tutorial is out of date.

judge.py
import pandas as pd
from phoenix.evals import LLM, ClassificationEvaluator, evaluate_dataframe

llm = LLM(provider="openai", model="gpt-4o-mini")

TEMPLATE = """You are checking a customer support answer.

Policy: {policy}
Question: {question}
Answer: {answer}

Is the answer fully consistent with the policy? Reply with exactly one word:
consistent or inconsistent."""

faithful = ClassificationEvaluator(
    name="policy_consistency",
    prompt_template=TEMPLATE,
    llm=llm,
    choices={"consistent": 1.0, "inconsistent": 0.0},
)

df = pd.DataFrame(
    {
        "policy": ["Refunds are available within 14 days of purchase with a receipt."] * 2,
        "question": ["Can I get a refund after two weeks?"] * 2,
        "answer": [
            "Refunds are only available within 14 days, so two weeks is the limit.",
            "Yes, refunds are available for up to 60 days.",
        ],
    }
)

results = evaluate_dataframe(dataframe=df, evaluators=[faithful])
print(results)

Read the template carefully. The placeholders in curly braces ({policy}, {question}, {answer}) must match the column names of the dataframe, and the last line tells the judge to reply with exactly one of the labels in choices. The choices dictionary maps each label to a numeric score. The first row should come back consistent with score 1.0 and the second inconsistent with 0.0. The exact layout of the results table depends on the library version, so print it and read it rather than assuming a column name.

Built-in evaluators

You do not always need to write templates. The library includes ready-made metrics under phoenix.evals.metrics, such as FaithfulnessEvaluator, CorrectnessEvaluator, CompletenessEvaluator, ConcisenessEvaluator, RefusalEvaluator, ToxicityEvaluator and RetrievalRelevanceEvaluator, plus the deterministic exact_match. Two deprecations to know: HallucinationEvaluator is deprecated in favour of FaithfulnessEvaluator, and the older document-relevance evaluators are deprecated in favour of RetrievalRelevanceEvaluator.

Evaluating real traces

The value appears when you score the traces already in Phoenix. The pattern is: fetch spans as a dataframe, run the evaluators over it, and log the results back as annotations so they appear in the interface.

score_traces.py
from phoenix.client import Client
from phoenix.evals.utils import to_annotation_dataframe

client = Client(base_url="http://localhost:6006")

spans = client.spans.get_spans_dataframe(project_identifier="support-bot", limit=100)
# ... run your evaluators over `spans`, producing `results` ...
client.spans.log_span_annotations_dataframe(
    dataframe=to_annotation_dataframe(results),
    annotation_name="policy_consistency",
    annotator_kind="LLM",
)

After this runs, each scored span shows a policy_consistency annotation in the interface, and you can filter on it. The middle step is elided because it depends on how your span columns map to your template placeholders; evaluators take an input_mapping for exactly that purpose, and Mid-level covers it in detail.

Trust, but verify the judge

An LLM judge is a model, and models are wrong sometimes. Before you believe a judge, compare it against the hand labels you made in the previous section. If you labelled thirty traces and the judge agrees on twenty-eight, it is useful. If it agrees on seventeen, fix the template before you chart anything. This is the most skipped step in evaluation work and the one that separates dashboards you can act on from dashboards that only look confident.

NOT_PARSABLE results If an evaluation returns NOT_PARSABLE, the judge's reply was either cut off by a token limit or did not match any of your labels. Raise the model's maximum output tokens, and make the template demand exactly one of the label words.
Try it
  1. Run judge.py and read the results for the two rows.
  2. Add a third row with an answer that is half right, and predict the label before you run it.
  3. Change the template so the judge must also explain itself, and compare.
the consistent and inconsistent labels you expected for the clear cases, and an interesting disagreement on the half-right answer. That disagreement is the conversation you should have with your judge prompt.

Datasets and experiments: turning problems into tests

You now can find bad answers. The next question is how to stop them coming back. The answer is a dataset and an experiment.

A dataset is a named collection of examples. Each example has an input, which is a dictionary of whatever your application takes, an optional expected output, which is also a dictionary, and optional metadata. Phoenix versions datasets: every insert, update or delete creates a new version, so you can always say exactly which examples an experiment ran against. Datasets are the same idea as a test suite, except the expected results are sometimes fuzzy.

There are three common sources. You can build one by hand from cases you care about, you can export rows from a dataframe, or, most usefully, you can select traces in the interface, such as all the ones marked thumbs_down, and add them to a dataset with a button. The second route in code looks like this:

make_dataset.py
from phoenix.client import Client

client = Client(base_url="http://localhost:6006")

dataset = client.datasets.create_dataset(
    name="refund-questions",
    inputs=[
        {"question": "Can I get a refund after two weeks?"},
        {"question": "Do I need a receipt to get my money back?"},
        {"question": "How long does shipping to Riyadh take?"},
    ],
    outputs=[
        {"answer": "No. Refunds are available only within 14 days."},
        {"answer": "Yes, a receipt is required."},
        {"answer": "Between 5 and 8 working days."},
    ],
)
print(dataset)

Each position in inputs pairs with the same position in outputs, so keep the lists the same length. In the interface, the Datasets page now lists refund-questions with three examples.

Experiments

An experiment runs a task, which is an ordinary Python function from an example to an answer, across every example in a dataset, then scores the outputs with evaluators. The task is usually your application's entry point, so the experiment measures the real behaviour of your code on a fixed set of cases.

run_experiment.py
from phoenix.client import Client
from phoenix.client.experiments import run_experiment

from support_bot import answer_question

client = Client(base_url="http://localhost:6006")
dataset = client.datasets.get_dataset(dataset="refund-questions")


def task(input):
    return answer_question(input["question"])


def mentions_14_days(output, expected):
    # a simple code evaluator: did the answer contain the key fact?
    return "14" in output if "14" in expected["answer"] else True


experiment = run_experiment(
    dataset=dataset,
    task=task,
    evaluators=[mentions_14_days],
)

The task function may declare any of the parameters input, expected, metadata or example, and Phoenix passes the matching pieces. Notice that experiments live in phoenix.client.experiments; the old phoenix.experiments module was removed in version 14, and the arguments are now keyword-only. After the run finishes, open the Experiments tab on the dataset. You will see one row per example with the output, the evaluator results, and a link to the full trace of that run. That link is the point: a failing example in an experiment is one click from the spans that explain it.

A useful extra is repetitions. Language models are not deterministic, so a case that passes once may fail on the next run. Passing repetitions=3 to run_experiment runs each example three times and exposes the flakiness.

The improvement loop in practice

Here is how these pieces combine into a routine a team can follow.

  1. You change your prompt or your retrieval logic.
  2. You run the experiment against the same dataset as last time.
  3. You compare the two experiment runs side by side. Did the pass rate rise? Did any example that used to pass now fail?
  4. If it improved and nothing regressed, you ship. If not, you open the failing example's trace and read the prompt.

This is the difference between "I think the new prompt is better" and "the new prompt passes 41 of 45 cases, up from 36, and breaks none". It also makes change safe. Nobody is afraid to modify the prompt when a fast check stands behind it.

Grow the dataset from failures Start small, with ten to twenty cases. Every time a user reports a bad answer, turn that trace into a dataset example before you fix it. After a month you have a test suite made entirely of problems that really happened, which is the most valuable kind.
Try it
  1. Create the refund-questions dataset and add two more examples of your own.
  2. Run the experiment, then open the Experiments tab and click through to the trace of one failing row.
  3. Edit the system prompt in support_bot.py, run the experiment again, and compare the two runs.
two experiment runs on the same dataset, each row linked to its trace, and a visible difference in the evaluator results. You have just done prompt iteration with evidence.

Prompts and the Playground

Until now the prompt lived inside your Python file. Phoenix also offers prompt management: prompts stored in Phoenix as named, versioned templates, together with the model settings to run them. Each save creates a new version, and you can attach tags such as production to a version, so the application can ask for "the prompt tagged production" instead of a hard-coded string. Changing the live prompt then becomes a tag move, with history, rather than a code deployment.

The Playground is the interface for editing and running a prompt. You pick a provider and model, write or paste a prompt with variables, and run it. Its powerful feature is that you can run the prompt over an entire dataset at once and see all the outputs in a grid, which makes it the quickest way to eyeball a change before you write any code. It can also start from a real span: open a trace, choose to replay that LLM call in the Playground, and you land in an editor holding the exact prompt that produced the bad answer, ready to tweak.

Playground needs a provider key. You can paste your key into the interface's provider settings, or administrators can store provider secrets under Settings → Secrets, where they are write-only. A provider that does not appear in the dropdown may simply be filtered by the PHOENIX_ALLOWED_PROVIDERS setting, which limits what the Playground shows.

For a first project you can ignore the programmatic side of prompts. Try the loop in the interface first: replay a bad span, edit the prompt, run it, and save a new version. Mid-level covers fetching a prompt version from code and formatting it for a provider call.

Try it
  1. Open one of your traces, find the LLM span and send it to the Playground.
  2. Change one instruction and run it. Compare the output with the original.
  3. Run the edited prompt over your refund-questions dataset and read all the outputs in the grid.
a Playground pre-filled with the real prompt, a changed reply after your edit, and a row of outputs across the dataset. Fast experiments without touching code.

Getting your data back out: client and command line

Sooner or later you want the data outside the browser, to chart it, to share it or to script a check. Phoenix gives you three routes.

The Python client

You have already met Client. A few calls cover most needs:

export.py
from phoenix.client import Client

client = Client(base_url="http://localhost:6006")

df = client.spans.get_spans_dataframe(project_identifier="support-bot", limit=200)
print(df[["name", "start_time", "status_code"]].head())

print(client.projects.list())
print(client.datasets.list())

get_spans_dataframe returns a pandas dataframe with one row per span and columns such as name, span_kind, start_time and status_code, with attributes as attributes.* columns. Treat the column names as something to print and check, because the exact set depends on what your spans contain. You can also pass a query to narrow it, for instance only root spans:

PYTHON
from phoenix.client.types.spans import SpanQuery

root_spans = client.spans.get_spans_dataframe(
    project_identifier="support-bot",
    query=SpanQuery().where("parent_id is None"),
)

The old helpers get_spans_dataframe on the top-level client and query_spans were replaced in version 14; if you see them, use the forms here.

The px command line

Phoenix also has a command line tool, published on npm as @arizeai/phoenix-cli. It needs Node.js 22 or newer. It installs a program named px:

BASH
npm install -g @arizeai/phoenix-cli
px setup
px trace list --limit 5
px span list --span-kind LLM --status-code ERROR
px dataset list

px setup is an interactive helper that writes a small file called .env.phoenix holding your endpoint and project, then waits for a trace to arrive so you know the connection works. px trace list prints recent traces in your terminal, which is handy over SSH or inside a coding agent where a browser is awkward. Configuration resolves in a fixed order: command-line flags win, then the process environment, then the active profile, then .env.phoenix, then the defaults. Knowing that order solves a whole class of "why is it talking to the wrong server" mysteries.

The assistant, briefly

Since version 17 Phoenix includes PXI, a built-in assistant you can chat with inside the interface, and in a terminal with the command pxi (which needs a server on version 20 or newer). It can help you explore traces and draft filters. It is optional, and an administrator can switch it off. Treat it as a convenience on top of the skills in this guide, not a substitute for learning to read a trace yourself.

Try it
  1. Run export.py and print the column names of the dataframe with print(df.columns.tolist()).
  2. Find the columns that hold the input and output text.
  3. If you have Node 22 or newer, install the CLI and run px trace list --limit 5.
a dataframe with one row per span, a clear sense of which columns hold your text, and, with the CLI, the same traces in your terminal. The data is yours to use anywhere.

Configuration: environment variables and the files behind them

Hard-coding endpoint="http://localhost:6006/v1/traces" in every script is fine for learning and awkward afterwards. Phoenix is designed to be configured with environment variables, named settings that live outside your code. You set them once per project, and the same code then works on your laptop and on a team server.

Client-side settings

These are the ones your application and tools read:

Variable What it does
PHOENIX_COLLECTOR_ENDPOINT Where traces are sent. Also a fallback base URL for clients.
PHOENIX_ENDPOINT Base URL for the Python client, the px command and MCP. Always a base URL. Python trace export does not read it.
PHOENIX_API_KEY Sent as a bearer token when the server has authentication on.
PHOENIX_PROJECT The project name. PHOENIX_PROJECT_NAME is an older alias; if both are set, PHOENIX_PROJECT wins.
PHOENIX_CLIENT_HEADERS Extra headers to send.
PHOENIX_GRPC_PORT Override the gRPC port, default 4317.

Two of these are easy to confuse. PHOENIX_COLLECTOR_ENDPOINT is for trace export. PHOENIX_ENDPOINT is for everything else that talks to the server's API. If you set only the second, your Python application's spans will go to the default location and you will wonder why the project you configured stays empty.

The .env.phoenix file

The SDKs and the command line automatically look for a file named .env.phoenix, walking up from the current directory. It holds the same variables:

.env.phoenix
PHOENIX_COLLECTOR_ENDPOINT=http://localhost:6006
PHOENIX_PROJECT=support-bot

Three rules. The name is .env.phoenix, not .env. The real process environment always beats the file. And if you ever put an API key in it, add the file to your .gitignore before the first commit, because a key pushed to a public repository is compromised the moment it lands. You can turn off file discovery with PHOENIX_DISCOVER_CONFIG=false.

The second rule bites people. If an old export PHOENIX_PROJECT=demo lives in your shell profile, it silently overrides every .env.phoenix you ever write. When traces land in the wrong project, run echo $PHOENIX_PROJECT and echo $PHOENIX_ENDPOINT first.

Server-side settings you will meet early

Variable What it does
PHOENIX_PORT and PHOENIX_GRPC_PORT The ports the server listens on, 6006 and 4317 by default.
PHOENIX_WORKING_DIR Where SQLite data lives. Defaults to .phoenix in your home folder.
PHOENIX_SQL_DATABASE_URL Use PostgreSQL instead of SQLite, for example postgresql://user:pass@host:5432/dbname.
PHOENIX_ENABLE_AUTH Turns on login. Off by default.
PHOENIX_SECRET Required once auth is on. At least 32 characters with a digit and a lowercase letter.

For PostgreSQL you need the extra driver: pip install 'arize-phoenix[pg]'. The server supports PostgreSQL 14 or newer. Use it for anything shared; SQLite is for development.

Turning on authentication

When you move beyond your own laptop, enable auth. Generate a secret and pass it in the environment:

BASH
export PHOENIX_ENABLE_AUTH=true
export PHOENIX_SECRET="$(openssl rand -hex 32)"
phoenix serve

The first login is admin@localhost with password admin. Change it immediately. After that, each application needs an API key, which you create in the interface under Settings, and then export as PHOENIX_API_KEY where your application runs. Enabling auth on a server that applications already use stops their traces until they hold keys, so plan the switch. Note also that since version 19 an API key can no longer create other API keys; keys are created by a person signed in to the interface.

If the server refuses to start after you turn auth on, read the message. It will say the secret must be set, or that it must be at least 32 characters long with a digit and a lowercase letter. The fix is a longer, better secret.

Data privacy and air-gapped machines

Phoenix stores prompts and completions, so think about what you send. Two settings are worth knowing. PHOENIX_TELEMETRY_ENABLED=false turns off anonymous web analytics from the user interface, and the documentation states that trace data is never collected. On a machine without internet access, the UI may wait many seconds trying to fetch web fonts; set PHOENIX_ALLOW_EXTERNAL_RESOURCES=false to stop that.

Retention is forever by default Phoenix keeps traces until you delete them. On a busy application the database grows without limit. Set PHOENIX_DEFAULT_RETENTION_POLICY_DAYS to a number of days, or configure a policy per project. Deletion cannot be undone, so choose the period with whoever owns the data rules at your company.
Try it
  1. Remove endpoint, protocol and project_name from your register() call, but keep auto_instrument=True.
  2. Create a .env.phoenix with the two variables above, and give register() the protocol "http/protobuf".
  3. Run the script and read the banner.
a banner showing the project and endpoint from the file, and a trace in the right project. Then run PHOENIX_PROJECT=other python yourscript.py and watch the environment variable beat the file.

Reading common errors

Most failures in the first week are one of a handful of patterns. The skill is reading the symptom and knowing where to look.

No traces appear, and no error

This is the most common problem by a wide margin, and the frustrating part is that nothing fails loudly. Work through it in order.

First, read the register() banner. It shows the endpoint and the transport. If it says gRPC and port 4317 while you started Phoenix in Docker without -p 4317:4317, nothing is listening, and spans are dropped. If you are behind an HTTPS load balancer that exposes only port 443, gRPC on 4317 will never arrive either.

Second, force the transport. Pass protocol="http/protobuf" and a full .../v1/traces endpoint, which is what this guide's examples do. There is one specific trap the documentation calls out: if an environment variable holds a URL ending in /v1/traces and you let Python infer the protocol, it rewrites the port to 4317 and keeps the path, and no spans arrive.

Third, check that you set PHOENIX_COLLECTOR_ENDPOINT, not only PHOENIX_ENDPOINT, and that no stale variable from your shell profile is overriding your file.

Fourth, if you use a short script, make sure it does not exit before spans are sent. With the default simple processor each span is sent as it finishes, so this is rare. With batch=True spans are buffered, and a script that ends within a moment may need to let the provider flush.

Warnings you will see

Message Meaning and fix
No OpenInference instrumentors found. Maybe you need to update your OpenInference version? Skipping auto-instrumentation. You set auto_instrument=True but no openinference-instrumentation-* package is installed. Install the one for your library, such as openinference-instrumentation-openai.
Could not infer collector endpoint protocol, defaulting to HTTP. The endpoint has neither a /v1/traces path nor the gRPC port. Pass protocol= explicitly.
WARNING: It is strongly advised to use a BatchSpanProcessor in production environments. You used the default batch=False. Fine while learning; set batch=True when the app serves real traffic.
Transient error StatusCode.UNAVAILABLE encountered while exporting traces to localhost:4317, retrying in ... The gRPC collector is unreachable. Same fixes as the no-traces checklist.

Errors from the server

Symptom Cause and fix
401 Invalid token or Expired token Auth is on and your key is missing, wrong or expired. Export PHOENIX_API_KEY.
403 on a write request The key belongs to a VIEWER, which is read-only. Use a member or admin key.
503 Server is at capacity and cannot process more requests Spans are arriving faster than the server can store them. Batch on the client, or sample.
415 Unsupported content type on /v1/traces Something sent JSON-encoded OTLP. Use the protobuf exporter, or gRPC.
400 The name 'note' is reserved... You named an annotation note. Rename it.

Errors from old tutorials

A very common source of confusion is code copied from a pre-2026 tutorial. These all point to the same cause: the API changed in version 13 or 14.

What you see What to do
ModuleNotFoundError: No module named 'phoenix.session.client', or an error on px.Client Use from phoenix.client import Client and Client(base_url=...).
ImportError for llm_classify, OpenAIModel or phoenix.experiments Use phoenix.evals.LLM with ClassificationEvaluator, and phoenix.client.experiments.
A 404 on /v1/evaluations Use the /v1/span_annotations, /v1/trace_annotations or /v1/document_annotations endpoints.
phoenix: error: when you put flags before serve Write phoenix serve --port ....
A startup error when PHOENIX_POSTGRES_HOST includes :5432 Put the port in PHOENIX_POSTGRES_PORT instead.

Permission denied on a mounted volume

If you run the -nonroot Docker image variant and mount a folder, the container user may lack write permission, and Phoenix reports "Permission denied" writing to disk. Give the volume the right ownership with chown, or use the regular image while you learn.

A five-line debug routine When traces are missing: read the banner, confirm curl /healthz works from the same machine as the app, force http/protobuf with the full /v1/traces URL, run echo $PHOENIX_COLLECTOR_ENDPOINT, and check the port in your Docker command. That sequence finds the cause in nearly every case.
Try it
  1. Deliberately break your setup: change the endpoint port to 6999 and run the script.
  2. Read the banner and the error output, and write down what they tell you.
  3. Fix it, run again, and confirm a new trace arrives.
a clear symptom from a wrong port, and confidence that you can recognise it. Breaking things on purpose, in a safe place, is the quickest way to learn an error message.

Putting it all together

Let us finish with one small project that uses everything above, end to end. The goal: a support bot for a fictional online shop, traced, reviewed, scored, tested and improved. It should take about an hour. Keep the server from the first section running.

Step one, set up. Create a fresh folder, a virtual environment, and install the packages:

BASH
python3 -m venv .venv && source .venv/bin/activate
pip install arize-phoenix arize-phoenix-otel arize-phoenix-client arize-phoenix-evals \
  openinference-instrumentation-openai openai pandas
phoenix serve

Run phoenix serve in one terminal, and do the rest in another with the same environment activated.

Step two, instrument. Reuse support_bot.py from earlier. Add using_attributes around a handle_message function with a session id and prompt_version set to "v1" in the metadata. Run it twelve times with a mix of refund, shipping and out-of-scope questions, using two or three session ids.

Step three, read. Open the project. Look at the tree of three or four traces. Open the Sessions view. Find a trace where the tool returned nothing. Write one sentence on why.

Step four, annotate. Hand-label at least eight traces as good or bad with a short explanation, using a name like quality.

Step five, build a dataset from the bad ones. Select the bad traces in the interface and add them to a dataset called support-regressions. Add a few good ones so the set is balanced.

Step six, evaluate. Write a judge like judge.py with a template suited to your shop. Run it over the dataset's inputs and outputs, and compare its labels against your hand labels. If it disagrees on several, adjust the template until it agrees on most.

Step seven, experiment. Run an experiment over support-regressions with your current code. Record the pass rate. Now improve answer_question so that an out-of-scope question gets an honest "I do not know" instead of a policy guess, change prompt_version to "v2", and run the experiment again.

Step eight, compare. Open the two experiment runs. Did the pass rate rise? Did any previously passing example fail? Open one of the failures and read the prompt.

Step nine, tidy. Move your endpoint and project into a .env.phoenix file, add it to .gitignore, set batch=True in register(), and write a two-paragraph README explaining how to start the server and run the experiment.

When you finish you have used every idea in this guide: the trace tree, manual and automatic instrumentation, sessions and metadata, human and automatic annotation, a dataset, an experiment and configuration by environment. The same loop is what professionals run on production systems, with larger data and more care.

What you can now do, and what comes next

You can start and check a Phoenix server, and you know which port carries what. You can trace an application automatically with register(auto_instrument=True) and explicitly with decorators, and you can read the resulting tree to find which step caused a bad answer. You can group traces into sessions, stamp them with users, versions and tags, and filter them. You can annotate by hand and from code, score traces with an LLM judge, and know why you must check the judge against human labels. You can turn failures into a dataset and use an experiment to compare two versions of your application. You can configure everything with environment variables and a .env.phoenix file, and you recognise the typical errors, especially the silent no-traces failure and the outdated-tutorial failures.

Just as important is what you now know to avoid: the old phoenix --port serve ordering, the removed px.Client, the removed llm_classify family, and the obsolete embedding and drift features. When a tutorial disagrees with this guide, check its date.

Here is where to go next.

  • Mid-level covers how tracing works under the hood, including processors, batching and sampling, richer context attributes, the evals library in depth with input mappings and the built-in metrics, the experiment API with repetitions, prompt management from code, and the px command line in full.
  • Senior covers running Phoenix as a platform: PostgreSQL, scaling, authentication and single sign-on, retention, upgrades, multi-team isolation, security and the point at which a different tool is the better choice.
  • For neighbouring tools in this catalogue, compare the approach with Langfuse, learn a model-judging library in RAGAS, and review Docker and Kubernetes for deploying the server itself.

If you keep one habit from this guide, make it this: before you change a prompt, open a trace and read what the model was actually given. Everything else is tooling around that simple act of looking.

Sources