Skip to content
Back to student guides
DeepEvalLLMOpsEvaluation & testing3 levels100 sectionsCovers DeepEval 4.2

The Complete DeepEval Guide

Unit-test LLM apps with DeepEval’s metrics, test cases and CI integration. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
18sections
36examples

This is part one of three. It covers everything you need to start testing an LLM application with DeepEval, not a teaser. By the end you can install the tool, pick a judge model, write a test case, score it with a metric, run it from the command line and from Python, load a dataset, read a failing result, and put the whole thing into a small project that a colleague can run with one command. Mid-level and Senior take the same topics further (tracing, component-level evals, CI gates, cost control, running DeepEval as a team platform), and nothing you learn here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas only stick once you have watched your own test pass, fail, and pass again for a reason you understand. The version this guide was checked against is DeepEval 4.2.7, released on PyPI on 29 September 2026. The project ships patch releases almost weekly, so when a command here behaves differently on your machine, run deepeval --version first and compare.

Why testing an LLM app is a different problem

Suppose you have built a chatbot that answers questions about your company's refund policy. You try five questions by hand, the answers look good, and you ship. Two weeks later someone edits the prompt to make the tone friendlier. Nobody re-tests the refund answers, because re-testing means reading a dozen paragraphs of prose and deciding whether each is still right. A week after that, a customer is told that refunds take 90 days, when the policy says 30.

Ordinary software avoids this with automated tests. A function that adds two numbers has one correct output, so the test compares add(2, 3) to 5 and moves on. An LLM application breaks that approach in three ways.

The output is free text. "Refunds are accepted within 30 days" and "You have a month to return it" mean the same thing and share almost no words. An equality check fails on a perfectly good answer.

The output changes from run to run. The same prompt can produce different wording on Monday and Tuesday. A test that passes once is not proof it will pass tomorrow.

"Correct" has many dimensions. An answer can be factually right but irrelevant to the question, relevant but invented from nothing, grounded in the documents but rude, or polite and accurate but three times longer than the user wanted. One boolean cannot hold all of that.

Teams first responded with what people now call "vibe checking": read some outputs, nod, ship. It works for a demo and fails the moment more than one person edits the system. The next response was to have a person grade outputs in a spreadsheet, which is accurate but slow and does not run on every pull request.

DeepEval is an open-source Python framework that automates that grading. Its own tagline is "Pytest for LLM apps", and the comparison is a fair one. You write test cases, you attach scoring functions called metrics, and you run the lot from a terminal or a CI pipeline. A failing metric fails the run, exactly like a failing assertion. The scoring is mostly done by asking a second language model (the "judge") to grade the first one, an approach usually called LLM-as-a-judge. DeepEval is built by a company called Confident AI, which also sells an optional hosted dashboard; the framework itself runs entirely on your machine and needs no account.

YOUR APPproduces an answer
→
TEST CASEinput + output + context
→
METRICa judge scores 0 to 1
→
PASS / FAILscore vs threshold

That diagram is the whole tool. Everything else in this guide is detail about one of those four boxes.

It is worth being honest about what this buys you and what it does not. A judge model is itself an LLM, so its verdicts are not infallible and they vary slightly between runs. DeepEval does not remove the need for human judgement; it moves the human from "read every output, every time" to "decide once what good looks like, then let the machine check it on every change". Teams that get value from it treat a metric the way they treat a smoke test: a cheap, fast signal that catches regressions, backed by occasional human review of a sample. You will see this attitude repeated in the Tips file and again at the Mid and Senior levels.

The neighbouring tools are worth knowing by name. If your interest is specifically retrieval-augmented generation, the RAGAS guide covers a sibling evaluation library with overlapping metrics. If you want to see production traces and scores over time, the Langfuse guide covers an observability platform that can receive evaluation scores. DeepEval is the piece that sits in your test suite and your pipeline.

Try it
  1. Pick any LLM feature you have used or built (a chatbot, a summariser, a classifier).
  2. Write down three questions you could ask it, and for each one, what a correct answer must contain.
  3. Now write down one way an answer could be correct yet still unacceptable (too long, rude, off-topic).
a list of three "must contain" rules and one extra quality rule. Those are exactly the things you will turn into test cases and metrics in the next sections.

The mental model: four nouns

DeepEval has a large API, but a beginner needs only four nouns. Learn these precisely and the documentation stops feeling overwhelming.

Test case. A test case is one recorded interaction with your application. In the single-turn form, which is the one we use throughout this guide, it is an LLMTestCase object. One input went in and one output came out. The only field you must supply is input; the others are optional and each metric tells you which ones it needs. The fields you will meet first are:

Field What it holds Who fills it
input The user's message only, not your system prompt You, from your dataset
actual_output What your application really answered Your app, at test time
expected_output The ideal answer, written by a person You, in advance
context Ground-truth background information for this input You, in advance
retrieval_context The chunks your retriever actually fetched this time Your app, at test time

The last two are easy to confuse, and the confusion causes real mistakes. context is what an ideal system should know: static text from your dataset. retrieval_context is what your pipeline really pulled from a vector database on this particular run, and it must be a list of strings, even when it holds one chunk. Passing a bare string raises a TypeError.

Metric. A metric is a scoring function. You give it a test case, and it returns a score between 0 and 1 and usually a human-readable reason. Each metric has a threshold, which defaults to 0.5, and the metric counts as passed when score >= threshold. Almost every built-in metric uses an LLM as the judge. Some, such as ExactMatchMetric, are plain deterministic code and need no model at all.

Dataset and golden. A golden is a template for a test case. It holds an input and usually an expected_output, but normally not an actual_output, because you have not run your app yet. A dataset (EvaluationDataset) is a collection of goldens. At test time you loop over the goldens, run your app on each one, and turn the result into a test case.

Test run. A test run is one execution of an evaluation over a group of test cases: a snapshot of how your app performed at one point in time. If you run it twice after a prompt change, you have two test runs to compare.

Put together:

Before the test
Golden
input + expected_output
At test time
Your app runs on the golden's input
LLMTestCase
adds actual_output
Scoring
Metric + judge LLM
score, reason, pass or fail

Everything runs in your own process. Only the judge call leaves your machine, and only to the provider you configure.

There are two more words you will hear and can safely ignore for now. A classifier is like a metric but returns a label from a fixed set (for example "refused" or "answered") instead of a number. A trace is a recording of the internal steps your app took, used for evaluating agents and individual components. Both are covered at the Mid level.

One more distinction matters on day one: reference-based versus referenceless metrics. A reference-based metric needs a ground truth to compare against, such as expected_output. A referenceless metric, such as AnswerRelevancyMetric, judges the output using only the input and the output itself. Referenceless metrics are easier to start with because you do not have to write ideal answers first, and they are the only kind you can run on live production traffic.

Try it
  1. Take the refund-policy example. Write one input, one expected_output, and one context sentence as plain text.
  2. Pretend the app answered "Refunds take 90 days". Write that as the actual_output.
  3. Decide which of your three texts a judge would need to notice the mistake.
you need expected_output or context to catch the error. The input and the wrong answer alone look like a perfectly fluent reply. That is why some metrics require more fields than others.

Installing DeepEval and checking the setup

DeepEval is a pure Python package with no native build step, so the install is identical on Linux, macOS and Windows. It needs Python 3.9 or later (and below 4.0). Always work inside a virtual environment, so that the dozens of dependencies it pulls in (Pytest, the OpenAI client, OpenTelemetry and others) stay out of your system Python.

On Linux or macOS:

BASH
mkdir deepeval-lab && cd deepeval-lab
python3 -m venv .venv && source .venv/bin/activate
pip install -U deepeval

On Windows PowerShell:

POWERSHELL
mkdir deepeval-lab; cd deepeval-lab
py -m venv .venv; .\.venv\Scripts\Activate.ps1
pip install -U deepeval

If activation is blocked on Windows by the execution policy, run Set-ExecutionPolicy -Scope CurrentUser RemoteSigned once and try again. Notice the prompt now begins with (.venv); that prefix tells you the environment is active. If you open a new terminal later, you must activate it again before deepeval will be found.

Verify the install with three commands:

BASH
deepeval --version        # prints the installed version, for example 4.2.7
deepeval --help           # lists every command
deepeval diagnose         # shows the effective configuration

deepeval --version is the quick sanity check. deepeval --help shows you the command surface; you will use only a few of them this week. deepeval diagnose is the one worth remembering: it prints the settings DeepEval is actually using, tells you which source won for each value (your shell, a .env file, or a default), and masks secrets so that only the last few characters of a key show. When something behaves oddly, deepeval diagnose is the first thing to run, and deepeval diagnose --json produces output that is safe to paste into a bug report.

There is an optional extra for a terminal viewer that you will want later:

BASH
pip install 'deepeval[inspect]'     # quotes are for zsh; on Windows use double quotes

It adds the Textual library that deepeval inspect needs. You do not need it yet.

Telemetry is on by default DeepEval sends anonymous usage events (event names, metric names, whether you run in Jupyter, a random anonymous ID, and a coarse region from your IP) through PostHog. It does not send your test data. If your employer or client requires no outbound analytics, set DEEPEVAL_TELEMETRY_OPT_OUT=1 in your environment before you run anything.

Two pieces of housekeeping will save you later. First, DeepEval writes a .deepeval/ folder into the directory you run it from, holding a cache and the latest results. Add it to .gitignore straight away. Second, DeepEval automatically loads a .env file from the current directory when you import it, and it never overrides variables already set in your shell. That is convenient and occasionally a surprise; the Configuration section explains the order.

BASH
printf '.venv/\n.deepeval/\n.env\n.env.local\n' >> .gitignore
Try it
  1. Create the folder and virtual environment above and install deepeval.
  2. Run deepeval --version and deepeval diagnose.
  3. In the diagnose output, find the line for the OpenAI key. It should be missing or masked.
a version number such as 4.2.7, and a diagnose report with no real key visible. You have a working install and have not yet configured a judge.

Your first test, with no API key

Before you touch a language model, prove the pipeline works using a deterministic metric. ExactMatchMetric compares the actual output to the expected output character for character. It calls no LLM, so it needs no key, costs nothing, and is perfectly repeatable. It is a poor measure of quality for free text, but it is an excellent way to learn the moving parts.

Create a file named test_smoke.py. The name matters: the test_ prefix is how Pytest, which DeepEval builds on, discovers tests.

test_smoke.py
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import ExactMatchMetric


def test_capital_of_france():
    test_case = LLMTestCase(
        input="What is the capital of France?",
        actual_output="Paris",
        expected_output="Paris",
    )
    assert_test(test_case, [ExactMatchMetric()])

Run it with the DeepEval command, not bare pytest:

BASH
deepeval test run test_smoke.py

deepeval test run wraps Pytest with the DeepEval plugin, which is what collects metric scores, prints the results table and writes the .deepeval/ files. You should see a passing test and a summary. Now break it on purpose: change expected_output to "Lyon" and run again. The run fails, exits with a non-zero status code (which is how a CI system knows to stop), and the assertion message looks like this:

TEXT
AssertionError: Metrics: Exact Match (score: 0.0, threshold: 1.0, strict: False, ...

Read it from left to right. The metric is called Exact Match; it scored 0.0; it needed at least 1.0; strict tells you whether strict mode, which forces pass or fail with no middle values, was on. Notice the threshold is 1.0 rather than the usual 0.5. That is because exact match is binary by nature, so the metric sets its own effective threshold. Most metrics do not do this, and their default stays at 0.5.

Three things happened that are worth naming. assert_test takes one test case and a list of metrics, runs each metric, and raises an AssertionError if any fails. deepeval test run is the runner that makes this a Pytest run with extra reporting. And because it is Pytest, everything you know about Pytest still applies: test functions are discovered by name, you can select one with deepeval test run test_smoke.py::test_capital_of_france, and fixtures and parametrisation work as usual.

Run through deepeval, not plain pytest Plain pytest will still execute your test function and the assertion will still fire, but you lose the DeepEval results table, the saved run files and the cloud upload. For every example in this guide, use deepeval test run.
Try it
  1. Save test_smoke.py and run deepeval test run test_smoke.py. Confirm it passes.
  2. Change the expected output to "Lyon" and run it again.
  3. Run echo $? (macOS/Linux) or echo $LASTEXITCODE (PowerShell) to see the exit code.
the first run exits with 0, the second with a non-zero number and an AssertionError naming Exact Match with score 0.0. That non-zero exit code is exactly what stops a bad change in CI.

Choosing and configuring a judge model

Nearly every interesting metric needs a judge: a language model that reads your test case and decides how good it is. DeepEval defaults to OpenAI, and the documentation shows the default model as gpt-5.4. Whichever provider you choose, it must be configured before you construct a metric that needs it, because DeepEval builds the judge at construction time.

The simplest route is an environment variable:

BASH
# macOS / Linux
export OPENAI_API_KEY="sk-..."
POWERSHELL
# Windows PowerShell
$env:OPENAI_API_KEY = "sk-..."

If you forget, you get this error the moment you create a judge-based metric, not later when you run it:

TEXT
deepeval.errors.DeepEvalError: OpenAI API key is not configured. Set OPENAI_API_KEY in your environment or pass `api_key` to OpenAIModel(...).

That message is precise: DeepEval reached for its default provider, found no key, and stopped. The fixes are to export the key, to switch provider, or to pass a different model= to the metric.

Switching provider is a one-line CLI command, and the documentation calls this the recommended way to do it, because the setting then applies to every metric in every test:

BASH
deepeval set-anthropic --model=claude-sonnet-4-5 --save=dotenv
deepeval set-ollama --model=deepseek-r1:1.5b
deepeval set-gemini --model=gemini-2.5-pro

Model names change quickly, so treat those as examples and use whichever model your account lists. The catalogue of providers includes OpenAI, Azure OpenAI, Anthropic, Amazon Bedrock, Gemini, Grok, DeepSeek, Moonshot, OpenRouter, LiteLLM, Portkey and local servers through Ollama. deepeval unset-<provider> reverses a choice. The --save=dotenv flag writes the settings into a .env.local file instead of only the current shell session; secrets are never written to DeepEval's own JSON settings file, which is why the .env.local entry in your .gitignore matters.

Why does the judge choice matter so much? Because the metric is only as reliable as the model grading it. A small local model may fail to produce the structured JSON DeepEval asks for, and you will see:

TEXT
Evaluation LLM outputted an invalid JSON. Please use a better evaluation model.

That is not a bug in your test. It means the judge cannot follow the output format. Use a stronger model, or read the documentation's guide on using custom LLMs.

There is also a privacy angle relevant to many employers in the Gulf and Egypt. The judge sees your test data: the user input, your app's output and any context. With a hosted provider, that text leaves your machine. If your data must stay inside a country or a cloud tenancy, choose a judge that runs there: Azure OpenAI or Bedrock in a region you control, or a local model through Ollama. This is a configuration decision, not a code change, which is one reason to set the provider through the CLI rather than hard-coding it.

Keep a cheap judge for development Every metric call costs tokens. While you are learning, point DeepEval at a smaller, cheaper model, and only move to your strongest judge for the runs that gate a release. Note that scores can shift a little between judges, so never compare a score from one judge with a score from another.
Try it
  1. Set your provider's key (or run deepeval set-ollama if you have Ollama installed).
  2. Run deepeval diagnose again.
  3. Confirm the key appears masked and the provider flag is on.
a diagnose report where the provider you chose is enabled and the secret shows only its last characters. If it shows nothing, your shell did not export the variable; set it again in the same terminal you use to run tests.

Your first LLM-judged test

Now the real thing. The official quickstart uses G-Eval, a metric where you describe in plain English what "good" means and let the judge apply it. It is the most flexible metric in the library, which makes it the right first one to learn.

test_correctness.py
from deepeval import assert_test
from deepeval.test_case import LLMTestCase, SingleTurnParams
from deepeval.metrics import GEval


def test_refund_answer_is_correct():
    correctness = GEval(
        name="Correctness",
        criteria="Determine whether the actual output is correct and consistent with the expected output.",
        evaluation_params=[
            SingleTurnParams.ACTUAL_OUTPUT,
            SingleTurnParams.EXPECTED_OUTPUT,
        ],
        threshold=0.5,
    )
    test_case = LLMTestCase(
        input="How long do I have to return a product?",
        actual_output="You can return items within 30 days of delivery.",
        expected_output="Customers have 30 days from delivery to request a refund.",
    )
    assert_test(test_case, [correctness])

Read it piece by piece, because every part is a decision you will make again.

name is a label that appears in the results table. criteria is your natural-language rule. evaluation_params lists which test-case fields the judge is allowed to look at; here the actual and expected outputs. This list is how you control what the judge sees, and it is also why the metric tells you when a field is missing. SingleTurnParams is the current name of this enum. Older tutorials use LLMTestCaseParams, which still imports but prints a DeprecationWarning saying it "is deprecated and will be removed in a future release". Always write SingleTurnParams.

Under the hood G-Eval works in two stages. First the judge turns your criteria into a short numbered list of evaluation steps. Then it scores the test case against those steps, and DeepEval weights the score by the judge's token probabilities so that you get a smooth number between 0 and 1 instead of a jagged 0 or 1. You can skip the first stage by supplying evaluation_steps yourself, which makes the metric more repeatable: you must provide either criteria or evaluation_steps, and if you provide neither, the error "Either 'criteria' or 'evaluation_steps' must be provided." appears when the metric is used.

Run it:

BASH
deepeval test run test_correctness.py

You should see the judge's score, often something like 0.8 to 1.0 for this pair, and a pass. Try changing the actual_output to "Returns are not accepted after purchase" and run again; the score collapses and the test fails. You have just written your first regression test for language behaviour.

Two details deserve attention. The score is not a probability that the answer is right; it is the judge's graded opinion against your criteria. And the threshold of 0.5 is DeepEval's default, not a law. A customer-support bot where wrong refund information has legal consequences might set 0.8; an internal brainstorming tool might accept 0.4. You choose it by looking at real scores, which the section on reading results explains.

A judge call is a paid, networked call Each assert_test with a judge metric makes several LLM requests. A G-Eval test makes at least two (steps, then the score). A test file with 50 test cases and three metrics can easily make several hundred calls. Start with five test cases, look at the bill, then scale.
Try it
  1. Save test_correctness.py and run it with deepeval test run.
  2. Note the score and read the reason if it is printed.
  3. Replace the actual output with a wrong claim and re-run.
a pass with a high score on the correct answer, and a failure with a low score on the wrong one. If both pass, your criteria is too vague; make it stricter, for example by naming the exact fact that must match.

Test cases in depth

You have used three fields so far. Here is how to think about all of them, because a test case is the unit of everything else.

input is what the user typed, and nothing more. A common beginner mistake is to paste the whole prompt, system instructions included, into input. Do not. If your app wraps the user's question in a long template, the template is configuration of your app, not part of the question, and putting it in input makes every metric judge your template instead of your behaviour. DeepEval's model is to keep the template separate and record it as a hyperparameter when you want to compare prompts (see the datasets section).

actual_output is what your application produced for that input. In a real test you do not type it; you call your app:

PYTHON
def answer(question: str) -> tuple[str, list[str]]:
    # your real code: retrieve chunks, call the model
    chunks = ["Refunds are available within 30 days of delivery."]
    reply = "You can return it within 30 days."
    return reply, chunks

and then build the test case from the result. The test case is a record of a run that already happened, which means a DeepEval test never calls your app "through" DeepEval; you call it, then hand DeepEval the evidence.

expected_output is what a perfect answer would say. It is optional, and a large share of useful metrics do not need it. Write it when you have a clear reference, for example a known-correct policy statement.

context and retrieval_context are the two RAG fields, introduced earlier. Use context when you have ideal ground-truth passages in your dataset. Use retrieval_context when you want to judge the retrieval step or check that the answer stays faithful to what was retrieved. A retrieval context is always a list:

PYTHON
from deepeval.test_case import LLMTestCase

test_case = LLMTestCase(
    input="How long do I have to return a product?",
    actual_output="You can return it within 30 days.",
    retrieval_context=["Refunds are available within 30 days of delivery."],
)

If you pass a string instead of a list you get TypeError: 'retrieval_context' must be None or a list of strings or RetrievedContextData. The fix is always square brackets.

Other optional fields exist for later: tools_called and expected_tools for agents, metadata and tags for filtering, token_cost and completion_time for cost reporting, and name for a readable label. You can ignore all of them for now.

Notice that no metric uses every field. Each metric page in the documentation has a "Required arguments" line. If a metric needs a field you did not supply, you get:

TEXT
deepeval.errors.MissingTestCaseParamsError: 'retrieval_context' cannot be None for the 'Faithfulness' metric

Read the message as a recipe: the metric is Faithfulness, the missing field is retrieval_context. Supply it. If you are running a large file where only some cases have that field, pass -s to deepeval test run, or set SKIP_DEEPEVAL_MISSING_PARAMS=1, and DeepEval will skip those metric and test-case pairs instead of failing.

Conversations are a separate shape. A multi-turn chatbot is represented by a ConversationalTestCase holding a list of Turn objects, each with a role of "user" or "assistant" and some content. The metrics for conversations are different (for example TurnRelevancyMetric, RoleAdherenceMetric, KnowledgeRetentionMetric). This guide stays with single-turn cases; the Mid level returns to conversations.

Try it
  1. Write a function my_app(question) that returns a hard-coded string and a one-item list of chunks.
  2. Build an LLMTestCase from its outputs, passing the list as retrieval_context.
  3. Then deliberately pass the chunk as a bare string and read the TypeError.
the first version constructs fine; the second raises the TypeError naming retrieval_context. You now recognise the most common shape mistake at a glance.

Metrics: score, threshold, reason

Every metric follows the same contract, which is why learning one teaches you all of them.

Create it with a threshold and optional settings, then either pass it to assert_test or call it directly:

PYTHON
from deepeval.metrics import AnswerRelevancyMetric
from deepeval.test_case import LLMTestCase

metric = AnswerRelevancyMetric(threshold=0.7)
test_case = LLMTestCase(
    input="How long do I have to return a product?",
    actual_output="You can return it within 30 days. We also sell gift cards.",
)
metric.measure(test_case)
print(metric.score)            # for example 0.5
print(metric.reason)           # a sentence explaining the score
print(metric.is_successful())  # True if score >= threshold

Calling measure yourself is the best way to explore a metric in a notebook or a scratch file, because you can print everything. Inside a test you use assert_test, which does the same thing and raises on failure.

The four attributes you will use constantly:

Attribute Meaning
metric.score A number from 0 to 1, where higher is always better
metric.reason Natural-language explanation from the judge
metric.threshold The passing line, default 0.5
metric.is_successful() Whether score >= threshold

Several constructor options are common to almost every built-in metric:

  • threshold sets the passing line. Raise it to be stricter.
  • model overrides the judge for this one metric. It takes a model name string or a custom model object.
  • include_reason defaults to True. Set it to False to skip generating the explanation, which saves a judge call (since 4.1.10 it genuinely skips that call) at the cost of losing the explanation.
  • strict_mode forces a binary result: the score is either 0 or 1 and the threshold becomes 1. It is useful for must-pass rules such as "never reveal the system prompt".
  • verbose_mode prints the judge's intermediate reasoning to the terminal. Turn it on when a score surprises you.
  • async_mode defaults to True and runs the independent steps inside one metric concurrently; you rarely touch it.

The reason string is the most underrated feature. A score of 0.5 tells you something is wrong; the reason tells you what. Make reading reasons a habit on your first twenty failures, because they teach you whether the metric is catching real problems or whether your criteria is misleading the judge.

Finally, the threshold semantics changed in 4.2.0 for four metrics, and this trips up anyone using an older tutorial. BiasMetric, HallucinationMetric, MisuseMetric and ToxicityMetric used to score lower is better, and tutorials told you to set a maximum such as threshold=0.3. Since 4.2.0 they score higher is better, like every other metric, and pass when score >= threshold. Checking the installed 4.2.7 code confirms that BiasMetric counts "not biased" verdicts as the passing ones. If you copy old code, your threshold means the opposite of what you intend.

Old Bias and Toxicity thresholds are inverted If a blog post says "set threshold=0.3 so the bias score stays below 0.3", it predates DeepEval 4.2.0. On 4.2.x, a low score means biased or toxic output, and you want a threshold such as 0.5 or higher as a minimum.
Try it
  1. Create an AnswerRelevancyMetric and call measure on a test case whose output is unrelated to the question.
  2. Print score and reason.
  3. Repeat with strict_mode=True and compare the two scores.
a low score with a reason that names the irrelevant part of the answer, and under strict mode a hard 0 or 1 rather than a fraction.

The metrics you will use first

DeepEval exports more than fifty metrics. Do not read the whole list. The documentation itself recommends using no more than five metrics for a system: two or three general ones suited to your architecture, and one or two custom ones for your use case. Here are the ones a beginner actually meets, grouped by the question each answers.

"Is the answer on topic?" AnswerRelevancyMetric scores how relevant actual_output is to input. It splits the answer into statements and asks the judge whether each one is relevant. It needs only input and actual_output, so it is the easiest first real metric.

"Did the answer stick to the facts I gave it?" FaithfulnessMetric checks whether the claims in actual_output are supported by retrieval_context. It is the standard check for hallucination in RAG systems. It needs input, actual_output and retrieval_context.

"Did retrieval fetch the right chunks?" Three metrics grade the retriever: ContextualPrecisionMetric (are the relevant chunks ranked above the irrelevant ones), ContextualRecallMetric (did retrieval fetch everything needed for the expected answer) and ContextualRelevancyMetric (how much of what was retrieved is relevant). They need retrieval_context, and the precision and recall ones also need expected_output.

"Does it meet my own definition of good?" GEval, which you have already used. Use it for tone, completeness, brand voice, format, anything you can describe in a sentence.

"Is it safe?" BiasMetric and ToxicityMetric look for biased or toxic language; PIILeakageMetric looks for personal data; NonAdviceMetric and MisuseMetric check that the app stays in its lane. Remember the 4.2.0 direction change described above.

"Is it exactly right?" ExactMatchMetric, PatternMatchMetric (a regular expression) and JsonCorrectnessMetric are deterministic and run with no judge. Whenever a rule can be written as code, prefer these: they are free, instant and never flaky.

Here is the RAG trio on one test case:

test_rag.py
from deepeval import assert_test
from deepeval.test_case import LLMTestCase
from deepeval.metrics import (
    AnswerRelevancyMetric,
    FaithfulnessMetric,
    ContextualRelevancyMetric,
)


def test_refund_rag():
    test_case = LLMTestCase(
        input="How long do I have to return a product?",
        actual_output="You can return it within 30 days of delivery.",
        retrieval_context=[
            "Refunds are available within 30 days of delivery.",
            "Shipping to the UAE takes 3 to 5 working days.",
        ],
    )
    assert_test(
        test_case,
        [
            AnswerRelevancyMetric(threshold=0.7),
            FaithfulnessMetric(threshold=0.7),
            ContextualRelevancyMetric(threshold=0.5),
        ],
    )

The interesting thing about this test is that it grades the whole pipeline in pieces. If FaithfulnessMetric fails while ContextualRelevancyMetric passes, the retriever did its job and the model invented something. If relevancy fails and faithfulness passes, the model honestly summarised the wrong documents. A single quality score cannot tell you which half to fix; three focused metrics can. Here the second chunk about shipping is irrelevant, so expect the contextual relevancy score to sit below 1.0 even though the test passes.

A focused metric set

  • Relevancy, faithfulness, one custom G-Eval
  • Each metric answers a separate question
  • A failure points at the component to fix
  • Three to five judge calls per case

Every metric switched on

  • Fifteen metrics "just in case"
  • Overlapping questions, contradictory signals
  • Nobody knows which failure to trust
  • Bills and runtimes several times higher
Try it
  1. Run test_rag.py as written.
  2. Change the actual_output to claim refunds take 90 days and run again.
  3. Note which of the three metrics fail and which still pass.
faithfulness fails, because 90 days contradicts the retrieved text, while answer relevancy probably still passes because the reply stays on topic. One change, one metric reacting: that is the diagnostic value of splitting quality into parts.

Writing your own criteria with G-Eval

Built-in metrics cover common ground. The metric you will write most often yourself is a G-Eval, because your product has requirements no library can guess: "answers must be in Modern Standard Arabic", "never promise a delivery date", "always mention the support hotline when apologising".

There are three ways to specify the rule, from loosest to tightest.

Criteria only. A sentence or two. The judge generates its own steps. Quick, and slightly variable between runs.

Evaluation steps. You write the steps yourself, as a list. The judge follows them literally, so scores are more repeatable and your reasoning is visible in code review.

Steps plus a rubric. You also define what score ranges mean. A rubric maps ranges between 0 and 10 to descriptions, which anchors the judge.

test_tone.py
from deepeval import assert_test
from deepeval.test_case import LLMTestCase, SingleTurnParams
from deepeval.metrics import GEval


def test_support_tone():
    tone = GEval(
        name="Support tone",
        evaluation_steps=[
            "Check that the actual output apologises when the input describes a problem.",
            "Check that the actual output gives a concrete next step.",
            "Penalise any blame directed at the customer.",
            "Penalise replies longer than four sentences.",
        ],
        evaluation_params=[
            SingleTurnParams.INPUT,
            SingleTurnParams.ACTUAL_OUTPUT,
        ],
        threshold=0.7,
    )
    test_case = LLMTestCase(
        input="My order arrived broken and nobody answers my emails.",
        actual_output=(
            "I am sorry your order arrived damaged. I have opened a replacement "
            "request, and you will get a confirmation email within one working day."
        ),
    )
    assert_test(test_case, [tone])

A few rules of thumb save hours. Write steps that a human reviewer could follow without asking you questions. Test one quality per G-Eval; a metric called "Good answer" that checks tone, length, accuracy and format produces a muddy number you cannot act on. Include in evaluation_params only the fields the rule needs. And prefer specific words ("names the support hotline") over abstract ones ("high quality").

G-Eval is for subjective or mixed criteria. When the rule is objective and can be expressed as a decision tree ("if the reply contains a price, check the currency is AED"), the DAG metric is more repeatable, and the Mid level covers it. Note that G-Eval is not affected by the eval-mode setting introduced in 4.2.x, which only changes the built-in QAG-style metrics; you can ignore that setting as a beginner.

Sanity-check a new G-Eval with a known-bad answer Whenever you write a custom metric, run it on two hand-written outputs: one you are sure is good and one you are sure is bad. If the scores do not separate clearly, fix the steps before you trust the metric on real data.
Try it
  1. Write a G-Eval for a rule from your own project, using evaluation_steps.
  2. Run it on one clearly good and one clearly bad output.
  3. Adjust the steps until the good output scores at least 0.3 higher than the bad one.
two scores with a visible gap. If the gap is small, your steps are too vague or test more than one quality at once.

Two ways to run: deepeval test run and evaluate()

DeepEval gives you two entry points, and choosing between them is mostly a question of where the code lives.

Inside a test file, through deepeval test run. This is the Pytest style you have used. Each test_ function calls assert_test for one test case. Use it when the evaluation belongs in your repository's test suite and in CI, because a failed assertion produces a non-zero exit code.

From a script or notebook, through evaluate(). This takes a whole list of test cases and a list of metrics and scores them all in one call, running them concurrently. It does not raise on failure by default; it prints a report and returns a result object. Use it for exploration, for notebooks, and for comparing prompts.

run_eval.py
from deepeval import evaluate
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric, FaithfulnessMetric

cases = [
    LLMTestCase(
        input="How long do I have to return a product?",
        actual_output="You can return it within 30 days of delivery.",
        retrieval_context=["Refunds are available within 30 days of delivery."],
    ),
    LLMTestCase(
        input="Do you ship to Egypt?",
        actual_output="We ship to Egypt in 4 to 7 working days.",
        retrieval_context=["We ship to the UAE, Saudi Arabia and Egypt."],
    ),
]

evaluate(
    test_cases=cases,
    metrics=[AnswerRelevancyMetric(), FaithfulnessMetric()],
)

Run it like any Python script with python run_eval.py. Because evaluate() runs test cases and metrics concurrently (up to 20 at once by default), it is much faster than the same work in sequential tests, and it is also the quickest way to hit your provider's rate limits, which the errors section addresses.

The deepeval test run command has flags worth learning early. They go after the file name:

Flag What it does
-v Verbose output, so you see metric detail
-x Stop at the first failure
-s Skip metrics whose required fields the test case lacks
-i Ignore errors instead of failing the run
-c Use the cache, reusing results for unchanged cases
-n 4 Run tests in four parallel processes
-d failing Display only failing cases in the results table
-id "prompt-v2" Attach a label to this test run

Anything after a bare -- is handed to Pytest itself, for example deepeval test run test_rag.py -- --tb=short.

The cache deserves a sentence. DeepEval stores results keyed by the test-case content plus the metric configuration. With -c, re-running an unchanged test costs no judge calls. That is wonderful while you are editing assertions around a test and dangerous if you forget it is on, because you may be looking at a stale score. If your results seem frozen after you changed something the cache does not track, run without -c.

A common question is whether to use one or the other. A sensible division for a beginner: use evaluate() while you are exploring, because the feedback loop is fast and you can print results; move stable checks into test_*.py files run by deepeval test run once you want them in CI.

Try it
  1. Run run_eval.py with python and read the printed table.
  2. Run deepeval test run test_rag.py -v and compare the amount of detail.
  3. Run the same test again with -c and watch it finish faster.
the script prints a summary of all cases at once; the verbose test run shows metric-level detail; the cached rerun makes no new judge calls and so completes in a fraction of the time.

Datasets and goldens: testing more than one case

One test case is a demo. A useful evaluation has dozens of cases that cover your real traffic. Hand-writing each LLMTestCase is tedious, and it also mixes two jobs: deciding what to ask, and running your app. DeepEval separates them with goldens.

A Golden holds the inputs you care about. An EvaluationDataset holds a list of them.

dataset_demo.py
from deepeval.dataset import EvaluationDataset, Golden

dataset = EvaluationDataset(
    goldens=[
        Golden(
            input="How long do I have to return a product?",
            expected_output="Customers have 30 days from delivery to request a refund.",
        ),
        Golden(
            input="Do you ship to Egypt?",
            expected_output="Yes, we ship to Egypt in 4 to 7 working days.",
        ),
    ]
)

You can also load goldens from files, which is how you will keep them once the list grows beyond a screen:

PYTHON
dataset = EvaluationDataset()
dataset.add_goldens_from_json_file(file_path="goldens.json")
# or: dataset.add_goldens_from_csv_file(...)

Check the docstring of the loader you use for the exact column or key names your version expects (help(dataset.add_goldens_from_json_file) prints it), because file formats are the kind of detail that shifts between releases.

Now combine a dataset with Pytest's parametrisation so that every golden becomes its own test:

test_dataset.py
import pytest
from deepeval import assert_test
from deepeval.dataset import EvaluationDataset, Golden
from deepeval.test_case import LLMTestCase
from deepeval.metrics import AnswerRelevancyMetric

from my_app import answer  # your real function: returns a string

dataset = EvaluationDataset(
    goldens=[
        Golden(input="How long do I have to return a product?"),
        Golden(input="Do you ship to Egypt?"),
        Golden(input="Can I change my delivery address?"),
    ]
)


@pytest.mark.parametrize("golden", dataset.goldens)
def test_answers_are_relevant(golden: Golden):
    test_case = LLMTestCase(
        input=golden.input,
        actual_output=answer(golden.input),
    )
    assert_test(test_case, [AnswerRelevancyMetric(threshold=0.7)])

The decorator makes Pytest create one test per golden, so a failure names the exact input that failed. This pattern, golden in, real app call, test case out, is the backbone of nearly every DeepEval suite you will see in the wild. The key discipline is that the golden is stable and version-controlled, while actual_output is regenerated each run.

You do not have to write all the goldens yourself. DeepEval includes a Synthesizer that generates goldens from your own documents, which is handy when you have a knowledge base and no test questions:

BASH
deepeval generate --method docs --variation single-turn \
  --documents ./policies/refunds.pdf --num-goldens 10 --output-dir ./synthetic_data

The same thing is available in Python through Synthesizer().generate_goldens_from_docs(document_paths=[...]). Synthetic goldens are a starting point, not a finished dataset. Read every one; some will be trivial, some off-topic, and the ones that matter most, the odd real questions your users ask, will only come from production logs and support tickets. A good habit is to add a golden every time a bug is reported.

If you want to compare prompt versions across runs, log what changed as hyperparameters so each test run records it:

PYTHON
from deepeval import evaluate

evaluate(
    test_cases=cases,
    metrics=[AnswerRelevancyMetric()],
    hyperparameters={"model": "gpt-5.4", "prompt_version": "v2", "temperature": 0.2},
)

The hyperparameters are stored with the run, so a month later you can still tell which configuration produced which scores.

Try it
  1. Create a dataset with five goldens taken from real questions in your project.
  2. Wire it into a parametrised test like the one above, with a stub answer() if you need one.
  3. Run deepeval test run test_dataset.py and read the per-golden results.
one test per golden in the report. When one fails, you can tell from the test id which input caused it, without rerunning anything.

Reading results and finding out why a test failed

A score is only useful if you can explain it. DeepEval gives you several ways to look inside.

The terminal table. After every run you get a summary listing each test case, each metric, its score, pass or fail, and often the reason. Use -d failing to hide passing rows once your suite is large.

Verbose mode. Adding verbose_mode=True to a metric prints the judge's intermediate work: the statements it extracted, the verdict on each and the final reasoning. This is how you answer "why did the judge say that?". Use it on a single failing case, not on a full suite.

PYTHON
from deepeval.metrics import FaithfulnessMetric

metric = FaithfulnessMetric(threshold=0.7, verbose_mode=True)

Result files. Each run writes JSON into the .deepeval/ folder, including .latest_test_run.json and a fuller .latest_run_full.json. They are the raw data for any script or tool that wants to post-process results.

The inspect viewer. If you installed deepeval[inspect], then deepeval inspect opens a terminal interface for browsing the latest run, filtering with /, and copying a case as JSON with y. It is the nicest way to explore a large run without leaving the terminal.

Confident AI. If you run deepeval login, your runs can be uploaded to the hosted dashboard, and deepeval view opens the latest one in your browser. This is optional. If you are on a corporate network or work with sensitive data, you can do everything in this guide without it.

When a test fails, work through a fixed routine. First, read the reason. Second, ask whether the failure is in your app or your test: is the actual_output really wrong, or is the expected output outdated? Third, if the verdict looks wrong, rerun that single case with verbose_mode=True to see the judge's reasoning. Fourth, if the judge is consistently mistaken, your criteria is probably ambiguous or your judge model too weak; fix those rather than lowering the threshold to make the test pass.

Setting thresholds is a craft. A method that works: run the suite once with a low threshold, collect the scores for outputs you know are good and outputs you know are bad, and set the threshold between the two clusters. Re-check the choice after you change judge models. Also expect variance. The same test case may score 0.82 on one run and 0.78 on the next. A threshold placed exactly at the typical score will flip between pass and fail, so leave a margin, and use the flaky=True option on borderline cases (available from 4.1.5) if you want a failing case to warn instead of failing the run.

Name your runs Pass -id "after-prompt-change" to deepeval test run, or identifier= to evaluate(). When you look at results a week later, a label you chose is worth far more than a timestamp.
Try it
  1. Take a failing test from earlier and add verbose_mode=True to its metric.
  2. Read the judge's intermediate output and locate the exact statement it objected to.
  3. Open .deepeval/.latest_test_run.json and find the same case and score.
you can point at the sentence in the answer that caused the drop, and you can find the same score in the saved file. Being able to explain a failure is the real skill this section teaches.

Configuration: keys, dotenv files and what to put where

Most configuration is environment variables, and DeepEval reads them in a fixed order. Understanding the order removes a whole category of "it works in my terminal but not in my editor" confusion.

When you import deepeval, it loads dotenv files from the current directory in this order: .env, then .env.<APP_ENV>, then .env.local, with later files winning. It never overrides variables that are already set in the process. So the effective order, strongest first, is: variables exported in your shell, then .env.local, then .env.<APP_ENV>, then .env, then DeepEval's saved non-secret settings, then built-in defaults. A stale key in an old .env is a classic cause of 401 Incorrect API key provided, and deepeval diagnose shows which source won for each value.

The settings a beginner needs:

Variable Purpose
OPENAI_API_KEY Key for the default judge
CONFIDENT_API_KEY Key for the optional Confident AI cloud (the old API_KEY name was removed in 3.9.9)
DEEPEVAL_TELEMETRY_OPT_OUT=1 Turn off anonymous usage events
DEEPEVAL_DISABLE_DOTENV=1 Stop automatic loading of .env files; set it before deepeval is imported
SKIP_DEEPEVAL_MISSING_PARAMS=1 Skip metrics whose required fields are missing
ENABLE_DEEPEVAL_CACHE Turn the results cache on

A minimal project layout puts the key in .env.local, which is git-ignored:

.env.local
OPENAI_API_KEY=sk-your-key-here
DEEPEVAL_TELEMETRY_OPT_OUT=1

Never commit that file. If a key does leak into a repository, revoke it at the provider immediately; deleting the commit is not enough, because the key stays in git history.

Two more configuration ideas are worth knowing now, and both come back later. Cost scales with the number of judge calls, so every include_reason=False, every -c cache hit, and every deterministic metric you substitute for a judged one lowers your bill. Pinning matters because the project ships frequently and prompts inside metrics change: the 4.1.10 release, for example, reworded metric prompts and noted borderline verdicts might shift slightly. Put an exact version in your requirements file (deepeval==4.2.7) and upgrade on purpose, not by accident.

requirements.txt
deepeval==4.2.7
pytest>=8
Try it
  1. Move your API key from your shell into a .env.local file and unset the shell variable.
  2. Run deepeval diagnose and find which file the key is read from.
  3. Confirm .env.local is listed in .gitignore.
diagnose names .env.local as the source of the key, and git status does not list the file as untracked. You now control where secrets live.

Common errors and how to read them

Most first-week failures fall into a handful of families. The pattern is always the same: read the exception class, read the quoted name inside it, and map it to the noun it concerns.

What you see What it means What to do
DeepEvalError: OpenAI API key is not configured... A judge metric was constructed with no provider Export OPENAI_API_KEY, run deepeval set-<provider>, or pass model=
openai.AuthenticationError ... invalid_api_key (401) The key is wrong or stale, often from an autoloaded .env Fix the key; run deepeval diagnose to find the winning source
MissingTestCaseParamsError: '<field>' cannot be None for the '<Metric>' metric The test case lacks a field the metric needs Fill the field, or run with -s to skip
TypeError: 'retrieval_context' must be None or a list of strings... You passed a string, not a list Write retrieval_context=["chunk"]
Evaluation LLM outputted an invalid JSON... The judge cannot follow the output format Use a stronger judge model
AssertionError: Metrics: <Name> (score: ..., threshold: ...) A metric failed. This is a normal test failure Read the reason; fix the app or the test
ValueError: No Confident API key found... A cloud action (push, pull, view, --official) without a key deepeval login, or skip the cloud step
DeprecationWarning: 'LLMTestCaseParams' is deprecated... Old import name Use SingleTurnParams
openai.RateLimitError (429), or a run that seems stuck Too many concurrent judge calls Lower concurrency (see below)
ModuleNotFoundError mentioning Textual on deepeval inspect The optional extra is missing pip install 'deepeval[inspect]'

Rate limits deserve a worked example, because every beginner meets them the first time they scale from three cases to three hundred. evaluate() runs up to 20 things at once by default. Your provider may allow fewer. Lower the concurrency with the async configuration:

PYTHON
from deepeval import evaluate
from deepeval.evaluate import AsyncConfig

evaluate(
    test_cases=cases,
    metrics=[AnswerRelevancyMetric()],
    async_config=AsyncConfig(run_async=True, max_concurrent=5, throttle_value=2),
)

max_concurrent=5 allows five at a time and throttle_value=2 adds a pause between them. Under deepeval test run, reduce the -n process count instead. By default DeepEval retries a failed call once, and a run on a slow network may also hit its time budget; both are tunable with environment variables that the Mid level covers.

A different kind of trap is a test that passes because it checks nothing. If you write a metric with a threshold of 0, or an assert_test with an empty metric list, the test passes forever. Likewise, a dataset loop that runs zero goldens passes silently. When you add a test, make it fail once on purpose to prove it can.

And a final, human error: trusting one run. If you change a prompt and one test that used to score 0.9 now scores 0.85, that is noise, not a regression. If twenty tests each drop by 0.1, that is signal. Look at averages across a dataset before you conclude anything.

Do not silence errors to get a green run Flags like -i (ignore errors) exist for flaky networks and weak local judges. If you use them to hide a real MissingTestCaseParamsError or an invalid JSON, the suite turns green while measuring nothing. Use them knowingly, and read what was ignored.
Try it
  1. Deliberately create each of three errors: no key, a string retrieval_context, and a Faithfulness metric with no retrieval context.
  2. Read each message and say aloud which noun it names.
  3. Fix each one.
three distinct messages that you can now map to "provider", "field shape" and "missing field" in seconds. Next time you meet them unprompted, you will not need a search engine.

Putting it all together

Here is one small end-to-end project that uses everything above: a refund-policy assistant with a dataset, three metrics, a deterministic check, and a pinned, reproducible setup. It has a simulated app so you can run it without any real retrieval system.

The layout:

TEXT
refund-evals/
  .gitignore
  requirements.txt
  .env.local            # not committed
  app.py
  goldens.py
  test_refund_bot.py

The application under test, which you would replace with your own code:

app.py
POLICY = {
    "return": "Refunds are available within 30 days of delivery.",
    "shipping": "We ship to the UAE, Saudi Arabia and Egypt in 3 to 7 working days.",
}


def retrieve(question: str) -> list[str]:
    q = question.lower()
    chunks = []
    if "return" in q or "refund" in q:
        chunks.append(POLICY["return"])
    if "ship" in q or "deliver" in q:
        chunks.append(POLICY["shipping"])
    return chunks or list(POLICY.values())


def answer(question: str) -> str:
    chunks = retrieve(question)
    # Replace this stub with a real LLM call that uses the chunks.
    return "According to our policy: " + " ".join(chunks)

The goldens, kept in their own file so they are easy to review:

goldens.py
from deepeval.dataset import EvaluationDataset, Golden

dataset = EvaluationDataset(
    goldens=[
        Golden(
            input="How long do I have to return a product?",
            expected_output="Customers have 30 days from delivery to return a product.",
        ),
        Golden(
            input="Do you ship to Egypt?",
            expected_output="Yes, shipping to Egypt takes 3 to 7 working days.",
        ),
        Golden(
            input="Can I pay with cash on delivery?",
            expected_output="The policy does not mention cash on delivery, so the bot should not promise it.",
        ),
    ]
)

And the tests, which combine judged metrics, a custom G-Eval and a deterministic rule:

test_refund_bot.py
import pytest
from deepeval import assert_test
from deepeval.test_case import LLMTestCase, SingleTurnParams
from deepeval.metrics import (
    AnswerRelevancyMetric,
    FaithfulnessMetric,
    GEval,
    PatternMatchMetric,
)

from app import answer, retrieve
from goldens import dataset


def make_metrics():
    correctness = GEval(
        name="Matches reference",
        evaluation_steps=[
            "Check that the actual output does not contradict the expected output.",
            "Penalise any promise that the expected output does not support.",
        ],
        evaluation_params=[
            SingleTurnParams.ACTUAL_OUTPUT,
            SingleTurnParams.EXPECTED_OUTPUT,
        ],
        threshold=0.6,
    )
    return [
        AnswerRelevancyMetric(threshold=0.7),
        FaithfulnessMetric(threshold=0.7),
        correctness,
    ]


@pytest.mark.parametrize("golden", dataset.goldens)
def test_refund_bot(golden):
    test_case = LLMTestCase(
        input=golden.input,
        actual_output=answer(golden.input),
        expected_output=golden.expected_output,
        retrieval_context=retrieve(golden.input),
    )
    assert_test(test_case, make_metrics())


def test_reply_never_claims_ninety_days():
    test_case = LLMTestCase(
        input="How long do I have to return a product?",
        actual_output=answer("How long do I have to return a product?"),
    )
    # A deterministic guard: fail if the forbidden number appears. No judge call.
    forbidden = PatternMatchMetric(pattern=r"^(?!.*\b90 days\b).*$")
    assert_test(test_case, [forbidden])

The last test uses PatternMatchMetric, a deterministic metric that takes a regular expression. The pattern here is a negative lookahead that matches only when "90 days" does not appear. Confirm it behaves on your machine with a quick failing case before you rely on it, since regular-expression flags such as whether the match must span the whole string are worth a five-minute check in your installed version.

Run the suite and tag the run:

BASH
pip install -r requirements.txt
deepeval test run test_refund_bot.py -id "baseline" -d failing

Read the results. The cash-on-delivery golden is the interesting one: a naive bot that answers "Yes, cash on delivery is available" would fail faithfulness, because nothing in the retrieved policy supports it, and would fail the reference check, because the expected output says the policy is silent. That is DeepEval doing its job. Then change app.py to hard-code a wrong answer, see a failure, and change it back.

To wire this into CI, the structure is the same as any Pytest job: install dependencies, export the judge key as a secret, and run deepeval test run. A GitHub Actions workflow would set OPENAI_API_KEY from secrets.* and run the command as a step; the GitHub Actions guide covers the workflow file. Set DEEPEVAL_DISABLE_DOTENV=1 in CI so that no stray .env file is loaded, and remember that a non-zero exit code is what makes the pipeline stop.

  1. Write the goldens. Start with five real questions and their expected answers.
  2. Run the app on each golden. Capture actual_output and retrieval_context.
  3. Attach three or four metrics. Mix judged ones with at least one deterministic guard.
  4. Run, read reasons, tune thresholds. Set them between your known-good and known-bad scores.
  5. Commit it and run it on every change. Pin the version and keep the key in a secret.
Try it
  1. Build the project above in a fresh folder and run it.
  2. Replace the stub in answer() with a call to your own model or application.
  3. Add two goldens that represent real failures you have seen, and confirm they fail before you fix them.
a green suite on the honest stub, a red test when you inject a wrong answer, and two real-world regression cases that now guard your project for good.

What you can now do, and what comes next

You started with a problem that has no natural assertion: free-text output that varies and has many dimensions of quality. You can now turn a behaviour into a LLMTestCase, score it with deterministic, RAG and custom metrics, run it through deepeval test run or evaluate(), feed it from a dataset of goldens, configure a judge provider without leaking a key, and read a failure down to the sentence that caused it. You also know the traps: thresholds that flipped in 4.2.0, the string versus list mistake, errors silenced to make a run green, and scores compared across different judges.

What comes next depends on your work.

  • Mid-level picks up with tracing and @observe so you can evaluate agents and individual components, conversation testing, the synthesizer and simulator in depth, the full deepeval test run flag set, caching and cost control, and a real CI workflow with baselines.
  • Senior covers running DeepEval as a platform for other teams: settings precedence, local storage, governance gates, judge-model security, multi-tenancy, upgrade discipline and the limits of LLM-as-a-judge.
  • Neighbouring tools. For dedicated retrieval-quality scoring, read the RAGAS guide. For production traces and dashboards that can hold your scores, read the Langfuse guide. For running evals on every pull request, read the GitHub Actions guide.

The single most valuable next step is unglamorous: take a real project, write ten goldens from real user questions, and run them. The first time a metric catches a regression you would have shipped, the tool has paid for itself, and you will know which metrics to trust in your own domain.

Sources