تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
HeliconeLLMOpsObservability & tracing3 مستويات120 قسمًايغطّي Helicone v2025.08 (maintenance mode)دليل بالإنجليزية

The Complete Helicone Guide

Log, monitor and control LLM usage and cost with the Helicone gateway. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
18sections
23examples

This is part one of three. It covers what you need to do real work with Helicone: send a language-model request through it, find that request on a dashboard, label your traffic so the numbers mean something, and switch on the gateway features that save money and absorb failures. By the end you will be able to explain where every request goes, read a cost figure without guessing, and recognise the handful of errors that catch almost every newcomer. Mid-level and Senior take the same topics further; nothing here is thrown away.

One honest note before we start, because it changes how you should use this guide. Helicone is in maintenance mode. On 3 March 2026 the company announced it was joining Mintlify, and on the same day new sign-ups on Helicone Cloud were switched off. The product still runs, security fixes and new models still ship, and no end date has been published, but you cannot simply create a fresh cloud account today. This guide is written for that reality: it teaches the ideas and the exact request format, tells you which parts you can practise on your own machine, and flags the places where older tutorials on the internet are now wrong. The skills transfer directly to any other LLM observability tool, which is part of why they are worth learning.

Each section ends with a Try it task. If you have a cloud account, do them against the cloud. If you do not, the self-hosting section shows you how to run the dashboard locally, and most tasks can be read as "what would I see" exercises. These ideas only stick once you have watched your own request appear, fail, and appear again properly.

What Helicone is, and where it stands today

Helicone is an open-source LLM observability platform with a gateway attached. Two halves of that sentence matter. Observability means that every call your application makes to a language model is recorded: what you sent, what came back, how many tokens each side used, what it cost, how long it took, and whether it succeeded. Gateway means that Helicone can sit in the path of those calls and do useful things to them on the way through, such as returning a stored answer instead of paying for a new one, retrying a failed call, or refusing a call that would exceed a limit you set.

YOUR APPcalls a model
→
HELICONElogs and adds features
→
MODEL PROVIDEROpenAI, Anthropic, ...
→
DASHBOARDcost, latency, requests

The picture is the whole product in four boxes. Your application talks to Helicone instead of talking straight to the model provider; Helicone forwards the call, hands you the answer, and records everything about the exchange so you can look at it later.

Now the status, stated plainly because it affects your plans:

  • 3 March 2026: Helicone announced it was joining Mintlify and would continue "in maintenance mode", meaning security updates, new models, and bug and performance fixes keep shipping, but the product is no longer growing new features.
  • Sign-ups are disabled. The sign-up page now says so. The official quick-start still says "sign up for free", and that line is stale. If you already have an account, it keeps working.
  • Self-hosting is open. The code is Apache-2.0 licensed on GitHub, and there is a Docker image you can run yourself. That is the path for a learner who has no cloud account.
  • Pricing is historical. The pricing page still lists a free tier of 10,000 requests a month and paid tiers, but self-serve upgrades have been removed, so treat those numbers as information rather than a shopping list.

There is no version number to install. Helicone is a continuously updated service plus a code repository. The newest tag on GitHub is v2025.08.21-1, and the self-host Docker image is helicone/helicone-all-in-one:v2025.08.21. Both are dated August 2025. When this guide says "the current docs", it means the documentation as checked in September 2026.

Why learn a tool in maintenance mode? Three reasons. First, plenty of teams run it in production today and will for a while. Second, the concepts (proxying, async logging, sessions, custom properties, cache keys, rate-limit windows) are the same ones every alternative uses, and Helicone's header-driven design makes them unusually easy to see. Third, you will be asked about it, or about moving off it, in interviews and in jobs. Knowing both the tool and its current status is exactly what a careful engineer is expected to know.

Do not trust old tutorials blindly Most blog posts about Helicone from 2023 to 2025 are out of date in at least one way that matters: they show provider-specific proxy URLs as the main route, they use a different model-name format, they describe Supabase-era self-hosting, or they teach an "Experiments" feature that was removed in August 2026. When a tutorial disagrees with this guide, check the official docs listed under Sources.
Try it
  1. Open the Helicone blog post announcing the Mintlify move and read it once.
  2. Write down, in one sentence each, what "maintenance mode" does promise and what it does not promise.
  3. Decide which setup you will use for the hands-on tasks: an existing cloud account, or a local self-hosted copy.
a clear note that maintenance mode promises fixes and new models but no new features and no end date, and that you know which path you are on before you write any code.

The problem it solves, and what came before

Calling a language model is easy. The first call takes four lines of code and works. The trouble starts a week later, when something happens that those four lines cannot answer.

Your bill is higher than expected, and you do not know which feature, which user, or which prompt caused it. A customer says "the assistant gave me a strange answer yesterday afternoon", and you cannot see what was actually sent to the model or what came back. The provider starts returning rate-limit errors in the middle of a demo. Latency creeps up and you cannot tell whether it is the model, your prompt, or your own code. A test run costs real money because it re-asks the same questions every time. Each of these is a question about a request that already happened, and a bare SDK call leaves no record.

Before tools like Helicone, teams solved this in three ways, each with a cost.

Print statements and log files. You write a wrapper that logs the prompt and response. This works until you need to sum cost by user, filter by error, or show a non-engineer what happened. You end up building a small analytics product by accident, and you must remember to add the wrapper to every call site.

The provider's own dashboard. OpenAI and Anthropic show you usage, but per API key and per day, not per user, per feature, or per conversation. You cannot see individual prompts in a way you can search, and if you use several providers you must check several dashboards.

General-purpose monitoring. Tools built for web servers record request counts and latencies, but they do not understand tokens, model names, or prompt cost. The interesting number, dollars per conversation, is something you have to calculate yourself.

Helicone's answer is to make the record automatic and to attach the labels you care about with almost no code. The core trick is unusual and worth noticing now: features are switched on by HTTP headers. A header is a small named value that travels with every web request. To cache responses you add Helicone-Cache-Enabled: true. To attribute the call to a user you add Helicone-User-Id: user-123. You do not install a framework or restructure your code; you add a line.

Here is how the old way and the Helicone way compare.

Without a gateway

  • Logging wrapper written and maintained by you
  • Cost worked out by hand from token counts
  • One dashboard per provider
  • Retries, caching and limits coded per call site
  • Past requests unsearchable after a bug report

With Helicone

  • Every request recorded automatically
  • Cost calculated from a maintained model price list
  • One dashboard across providers
  • Retries, caching and limits switched on with headers
  • Filter by user, session, property, status or model

An honest caveat belongs here as well. Putting a gateway in front of your model calls adds one more thing that can fail, and you are trusting a third party with your prompts. We will return to both points, because choosing between the two integration styles in the next sections is exactly a choice about that trade-off.

Try it
  1. Think of an LLM feature you have built or used. List three questions you would want answered a week after launch.
  2. For each, note whether the provider dashboard could answer it.
  3. Circle the ones that need per-user, per-feature, or per-conversation detail.
most of your questions are circled. Those circled questions are the job Helicone does, and the rest of this guide teaches the labels that answer them.

The mental model: a few nouns

Helicone has a small vocabulary. Learn these words and the dashboard, the docs and the error messages all become readable.

Request. One logged call to a model. Every request gets a unique ID called the Helicone-Id, a UUID such as 3f2b8c1e-.... It comes back to you in a response header named helicone-id, and you can also choose your own ID by sending Helicone-Request-Id. The ID is how you attach a rating or a score to that exact call later.

Helicone API key. A secret that identifies your organisation to Helicone. It starts with sk-helicone-. You create it in the dashboard under Settings and API Keys. This is not your OpenAI or Anthropic key, and confusing the two is the single most common beginner mistake, so hold on to the distinction.

Provider key. Your key for the model provider itself, such as an OpenAI key. With the current gateway you store it inside Helicone, in a settings page called Provider Settings, and Helicone uses it on your behalf. The docs call this BYOK, "bring your own key". You are billed by the provider, not by Helicone.

AI Gateway. Helicone's hosted endpoint at https://ai-gateway.helicone.ai. It speaks the OpenAI request format, so the standard OpenAI SDK works against it, and it can translate that format to other providers behind the scenes. You authenticate to it with your Helicone key alone.

Legacy proxy. The older way: provider-specific addresses such as https://oai.helicone.ai/v1 for OpenAI and https://anthropic.helicone.ai for Anthropic. Here you send your provider key in the usual place and add your Helicone key in a separate header. The docs now describe these as "maintained but no longer actively developed". You will meet them in existing code, which is why we cover them, but new work should start on the AI Gateway.

User. The person using your application, identified with the Helicone-User-Id header. It is whatever string you choose. It powers per-user cost and per-user rate limits.

Custom property. Any label you invent, set with a header named Helicone-Property- followed by a name. Helicone-Property-Environment: production creates a property called Environment. You can filter and group by it.

Session. A group of related requests that belong to one task, such as a conversation or one run of an agent. You group them with Helicone-Session-Id, and you can give each step a place in a tree with Helicone-Session-Path.

Cache. Stored responses that Helicone can return for a repeated identical request, so you do not pay the provider twice.

The next diagram shows how the labels ride along with one call.

Your code
OpenAI SDK, base URL set to the gateway
request + Helicone headers
Helicone gateway
checks cache, limits, retries
forwards using your provider key
Model provider
returns the completion
response, then an asynchronous log entry
Dashboard and API
filter by user, session, property

There is one more idea that is easy to miss and important to get right: logging is asynchronous. Helicone returns your response first and records the log afterwards, through an internal queue. That design means a slow dashboard should not slow your users, and it is why the May 2026 incident, in which logging and the dashboard were down for about four days while the gateway kept serving requests, did not break customer apps. The logs were queued and not lost.

All header values are strings Every Helicone header takes a string. Write "true", not true, and "3", not 3. A numeric or boolean value in the wrong type is silently ignored by some client libraries, and the feature simply does not turn on. This single rule explains a large share of "my header does nothing" reports.
Try it
  1. Without looking back, write the difference between a Helicone key and a provider key in two sentences.
  2. For the AI Gateway and for the legacy proxy, write which key goes in the Authorization header.
  3. Check your answers against the text above.
on the AI Gateway only the Helicone key is sent, in Authorization. On the legacy proxy the provider key goes in Authorization and the Helicone key goes in a separate Helicone-Auth header.

Two ways in: through the gateway, or beside it

Before any code, choose how Helicone connects to your application. There are exactly two styles, and everything else in the docs is a variation.

Proxy, or gateway, integration. Your application sends its model calls through Helicone. Helicone is in the path: request in, forward to the provider, response out. Because Helicone sees the request before the provider does, it can do things to it: return a cached answer, retry a failure, enforce a rate limit, block a risky prompt. You get every feature. The price is that Helicone is now on your critical path. If it were unreachable, your calls would fail unless you added your own fallback.

Async integration. Your application calls the provider directly, exactly as before, and afterwards sends a copy of the exchange to Helicone for recording. Helicone is not in the path, so it cannot slow or break your calls. The price is capability: with async logging you get observability only, with no caching, no rate limits, and no gateway retries, because those all need to happen before the provider is called.

Proxy or gateway

  • Caching, retries, rate limits, routing
  • Change one base URL, add headers
  • Helicone is on your critical path
  • Provider traffic goes via Helicone

Async logging

  • Logs and metrics only
  • Install a small logger library
  • Helicone is off your critical path
  • Provider traffic stays direct

A rule of thumb that works for a beginner: start with the gateway because it is one URL change and you get everything; move to async if your organisation insists that no third party may sit in the request path, or if you need to log calls from a model that the gateway does not reach.

Two smaller details. The async route comes in flavours: the OpenLLMetry loggers (helicone-async for Python, @helicone/async for Node) wrap the provider SDKs automatically, and the manual logger (helicone-helpers, @helicone/helpers) lets you record a call by hand, which is useful for models the automatic wrappers do not cover. And the gateway comes in two generations, the current AI Gateway and the legacy provider proxies. Everything in this guide defaults to the AI Gateway.

Ask the availability question early The useful question is not "which is better" but "if the observability tool disappears for an hour, should my product still answer users?" If yes, you want async logging, or a gateway with your own fallback path. If a short outage is acceptable, the gateway is simpler. Say this out loud in design discussions; it is exactly what a senior reviewer will ask.
Try it
  1. For each of these, write "gateway" or "async": you want to cache repeated answers; you must not add a network hop; you want a per-user spending cap.
  2. Explain in one sentence why the answer for each follows from where Helicone sits.
caching and per-user caps need the gateway, because they act before the provider is called. "No extra network hop" points to async.

Getting set up: which path fits you

There are three setups, and only you know which applies. All three end the same way: you have a Helicone API key and a place where requests show up.

Path A: you already have a cloud account

Log in at https://us.helicone.ai, or https://eu.helicone.ai if your organisation was created in the EU region. Open Settings, then API Keys, and generate a key. It begins with sk-helicone-. Put it in an environment variable rather than in your code:

BASH
export HELICONE_API_KEY="sk-helicone-xxxxxxxx-xxxxxxx-xxxxxxx-xxxxxxx"

On Windows PowerShell the equivalent is $env:HELICONE_API_KEY="sk-helicone-...", and in the old command prompt it is set HELICONE_API_KEY=sk-helicone-.... The variable lasts for that terminal window only.

Then add your provider key inside Helicone, at the Providers page (https://us.helicone.ai/providers). This step matters more than it looks. The docs' headline flow says you can "add credits" and skip provider keys, but that billing route (called pass-through billing) now requires a feature flag that most accounts do not have, since an August 2026 change. Without the flag it is denied. Bringing your own provider key works for everyone, so make it your default.

Finally install only the OpenAI SDK, since the gateway speaks its format:

BASH
pip install openai          # Python
npm install openai          # Node

Path B: you have no account, and want to learn

Run Helicone on your own machine. This is the route for a newcomer, and the section named "Running Helicone yourself" walks through it. Be aware of a documented limitation there: the local image gives you the dashboard and the log store but its documentation does not currently describe how to route model traffic through it. We will be precise about what you can and cannot do.

Path C: you are reading along only

That is fine. Every example in this guide is complete, and the expected results are described. Treat the Try it tasks as thought exercises.

Checking the setup

Whatever your path, the test of a working setup is the same. A request is sent, and it appears on the Requests page of the dashboard, which the docs say happens "within seconds". You can also check the response headers: a healthy call carries a helicone-id header. The next section does exactly this.

Never paste keys into code you commit Both keys are secrets. Read them from environment variables, keep any .env file out of git by adding it to .gitignore, and if a key ever lands in a repository, revoke it in the dashboard and create a new one. Deleting the commit does not help, because the history keeps it.
Try it
  1. Set HELICONE_API_KEY in your terminal using the form for your system.
  2. Run echo $HELICONE_API_KEY (PowerShell: $env:HELICONE_API_KEY) and confirm the value starts with sk-helicone-.
  3. Close the terminal, open a new one, and run the same check again.
the key prints the first time and is empty the second. That is what "set for this terminal only" means, and it is the most common reason a script works on Monday and fails on Tuesday.

Your first request

Everything starts with one call. We use the AI Gateway and the standard OpenAI SDK, changing exactly two things from a normal OpenAI program: the base URL, and the key.

first_request.py
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://ai-gateway.helicone.ai",
    api_key=os.getenv("HELICONE_API_KEY"),
)

response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Hello, world!"}],
)

print(response.choices[0].message.content)

Read it line by line, because each line is a lesson.

base_url tells the OpenAI SDK where to send requests. Normally it points at OpenAI's servers; here it points at Helicone's gateway. Nothing else in your program needs to know that anything changed. api_key is your Helicone key, not an OpenAI key, because the gateway authenticates you and uses the provider key you stored earlier to talk to the model. The model string is a plain model name. On the current gateway, a bare name like gpt-4o-mini lets Helicone pick a provider that offers it, cheapest first, and fail over to another if one is down.

The same call from the command line, which is the best way to see exactly what travels over the wire:

BASH
curl https://ai-gateway.helicone.ai/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $HELICONE_API_KEY" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Hello, world!"}]}'

And the Node version:

first_request.ts
import { OpenAI } from "openai";

const client = new OpenAI({
  baseURL: "https://ai-gateway.helicone.ai",
  apiKey: process.env.HELICONE_API_KEY,
});

const response = await client.chat.completions.create({
  model: "gpt-4o-mini",
  messages: [{ role: "user", content: "Hello, world!" }],
});

console.log(response.choices[0].message.content);

The gateway exposes a few endpoints you will use: POST /chat/completions for chat, POST /responses for OpenAI's newer Responses format, and GET /v1/models to list what you can route to. Some pages in the docs show a /v1 prefix on these paths. The quick-start uses the root without it, and this guide does the same.

Now verify. Open the Requests page in your dashboard. You should see one row: model, status 200, a few tokens, a cost measured in fractions of a cent, and a latency. Click it to see the full request and response. That single view, every prompt and every answer with its cost, is the thing you did not have before.

To see the request ID from code, ask curl to show headers:

BASH
curl -i https://ai-gateway.helicone.ai/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $HELICONE_API_KEY" \
  -d '{"model":"gpt-4o-mini","messages":[{"role":"user","content":"Say hi"}]}' | head -20

Look for a line beginning helicone-id: in the response headers. Save it: later you will attach a score to that exact call.

The two most common first-call failures A 401 with an "invalid API key" style message usually means a key problem: you pasted the provider key where the Helicone key belongs, or the provider key stored in Helicone is wrong. A 429 saying "Insufficient credits" usually means you have no provider key stored and your account is falling back to the gated credits route. In both cases the first place to look is the Providers page, not your code.

If you are maintaining older code, you will see the legacy form. It sends the provider key normally and adds the Helicone key as a header:

legacy_openai.py
import os
from openai import OpenAI

client = OpenAI(
    api_key=os.environ["OPENAI_API_KEY"],
    base_url="https://oai.helicone.ai/v1",
    default_headers={"Helicone-Auth": f"Bearer {os.environ['HELICONE_API_KEY']}"},
)

Note the Bearer prefix inside the Helicone-Auth value. Forgetting it is a classic source of 401 errors on the legacy proxy. Anthropic's equivalent uses base_url="https://anthropic.helicone.ai" with the same header. Recognise this shape, keep it working if you inherit it, but write new code against the AI Gateway.

Try it
  1. Run the Python or curl example above with your own key.
  2. Open the Requests page and find your request. Click into it.
  3. Write down the model, token counts, cost and latency you see.
  4. Change the prompt, run it again, and watch a second row appear.
two rows, each with prompt and response bodies you can read. If a row does not appear, re-check which key you used and whether the request actually returned 200 in your terminal.

Reading the dashboard

A dashboard is only useful if you know what each number means, so spend a few minutes learning to read this one. Three areas matter for a beginner.

The Requests page. One row per call. The columns you will look at constantly are the timestamp, the model, the status code, the token counts (prompt tokens are what you sent, completion tokens are what came back), the latency, and the cost. You can filter by any of them, and by the labels you attach later, such as user and custom properties. Clicking a row opens the full request and response bodies, which is how you answer "what exactly did the model see?".

Cost. Helicone calculates cost from the usage numbers the provider returns and from its own price list, an open-source registry covering 300+ models. The figure is an estimate from public prices, not your invoice. If you negotiated pricing or use a model the registry cannot resolve, it may differ or be missing. Treat it as accurate to a few percent for ordinary models and verify large numbers against the provider's bill.

Latency. The time from your request arriving to the response finishing. For streaming responses there is also a time to first token, the delay before the first words appear, which is what users actually feel. A rising latency line can mean a slower model, longer prompts, or provider trouble, and filtering by model tells you which.

A token, for readers new to language models, is a chunk of text of roughly three or four characters. Models read and write tokens, and providers charge per token, with output tokens usually costing more than input tokens. This is why long prompts and long answers are both expensive, and why the dashboard shows the two counts separately.

Here is a realistic record to practise reading. Imagine the row says: model gpt-4o-mini, status 200, prompt tokens 42, completion tokens 18, latency 820 ms, cost $0.00001. You can tell at a glance that the call succeeded, was short, and cost a fraction of a cent. If another row says status 429 with 0 completion tokens, the call failed before generating anything, and the body will contain the provider's error message explaining why.

Status codes you will see 200 is success. 400 means the request itself was malformed or, with moderation on, flagged. 401 is an authentication problem. 429 is "too many requests" or "no credit" and can come from the provider or from a Helicone limit you set. 500 and above are provider or gateway faults. Reading the status first tells you whether to look at your code or at the provider.
Try it
  1. Send three requests with very different lengths: one short question, one asking for a long essay, one with a long pasted paragraph.
  2. On the Requests page, sort or scan by cost and by latency.
  3. Note which of the three is the most expensive and which is the slowest, and whether they are the same one.
the essay request shows the most completion tokens and the highest latency; the pasted paragraph shows the most prompt tokens. Cost follows token counts, which is the whole reason the dashboard splits them.

Naming your traffic: users, properties and sessions

A list of requests is useful; a list you can slice is powerful. Labels are how you slice, and all of them are headers you add to a call. This is where Helicone starts paying for itself.

Setting headers

There are two places to set headers. Per request, when one call needs a specific label:

per_request_headers.py
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Summarise our refund policy."}],
    extra_headers={
        "Helicone-User-Id": "user-123",
        "Helicone-Property-Feature": "refund-bot",
        "Helicone-Property-Environment": "production",
    },
)

Or once for the whole client, so every call carries them:

default_headers.py
client = OpenAI(
    base_url="https://ai-gateway.helicone.ai",
    api_key=os.getenv("HELICONE_API_KEY"),
    default_headers={
        "Helicone-Property-Environment": "production",
        "Helicone-Property-App": "support-assistant",
    },
)

In Node the per-request form passes a second argument: client.chat.completions.create({...}, { headers: { "Helicone-User-Id": "user-123" } }), and the per-client form is defaultHeaders.

Users

Helicone-User-Id is the identifier of the person using your application. Use a stable, non-sensitive ID such as your database's user ID, not an email address or a name, because these values are stored and shown in the dashboard. With it you can see cost per user, find your heaviest users, and later limit spending per user. Note that the per-user rate limit needs this header to exist; without it there is nobody to count against.

Custom properties

A property is any label you choose. The header is Helicone-Property- followed by a name, and the value is free text. Good properties answer the questions you listed earlier: Environment (production, staging, local), Feature (which part of your product made the call), PromptVersion, Tenant for a multi-customer product. Once set, you filter and group by them on the dashboard, set alerts on them, and even rate-limit by them.

Choose property names with care. They are the columns of your future analysis, so use consistent capitalisation and a small, agreed set across the team. Environment and environment would be two separate properties.

Sessions

Many modern applications make several model calls to answer one user request: an agent plans, searches, reads and writes. A session groups those calls so you can see the whole story. Three headers work together:

Header Example What it does
Helicone-Session-Id a UUID Groups all calls of one run or conversation
Helicone-Session-Path /plan or /research/search Places this call in a tree, so a call inside a step nests under it
Helicone-Session-Name Customer Support Groups sessions of the same kind so you can compare them
sessions.py
import uuid

session_id = str(uuid.uuid4())

def ask(question, path):
    return client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": question}],
        extra_headers={
            "Helicone-Session-Id": session_id,
            "Helicone-Session-Path": path,
            "Helicone-Session-Name": "Trip Planner",
        },
    )

ask("List three cities worth visiting in Jordan.", "/plan")
ask("What is the best month to visit Petra?", "/plan/timing")
ask("Write a three-day itinerary for Amman.", "/write")

Paths start with a slash and should describe what the step is, not when it ran: group by function (/research, /write), not by time. The docs list all three session headers as required together, so send all three for a proper tree. On the Sessions page you then see one session with three requests, the total cost of the run, and the tree shape, which is how you find the one expensive step in a ten-step agent.

Decide your labels before you need them You can only slice by labels that were attached when the request happened. A request logged without a Feature property stays without one forever. Spend ten minutes choosing three or four properties on day one (environment, feature, tenant, prompt version) and add them as client defaults. Retroactive labelling is possible through an API for a single request, but not something you want to do for thousands.
Try it
  1. Add Helicone-User-Id and two properties (Environment and Feature) to your client defaults.
  2. Run the three-call session example above.
  3. On the dashboard, filter the Requests page by Feature, then open the Sessions page and find "Trip Planner".
your filter narrows the list to your labelled calls, and the session shows three requests with a combined cost. If the session is missing, check that all three session headers were sent and that their values are strings.

Provider keys, model names and routing

The AI Gateway decides which provider serves your request from the model string, so it is worth understanding its small grammar. You will not need every form on day one, but you should recognise them when you read other people's code.

You write What happens
gpt-4o-mini Helicone considers every provider offering this model, tries the cheapest first, balances equal prices, and fails over if one errors
gpt-4o-mini/openai Pinned to OpenAI, with no failover
gpt-4o-mini/azure,gpt-4o-mini/openai,gpt-4o-mini An ordered chain: try Azure, then OpenAI, then anything else
!openai,gpt-4o-mini Any provider except OpenAI

The suffix after the slash names the provider. This is the current form. An older README for Helicone's standalone Rust gateway used a prefix (openai/gpt-4o-mini) and a different base URL ending in /ai. That project has been abandoned, so if you see the prefix form in a tutorial, it is from a previous generation, and mixing the two forms will fail.

Why does routing matter to a beginner? Because it explains two behaviours you will otherwise find mysterious. First, the same model name can be served by different providers on different calls, which shows in the dashboard's provider column. Second, automatic failover means a provider outage may be invisible to your users, at the cost of your calls sometimes landing on a provider you did not expect. If that matters for compliance or data residency, pin the provider with the suffix.

On the question of whose key is used: the docs say the gateway tries your own keys before Helicone-managed credits, but a second docs page says the opposite order. The two pages disagree, so do not build a design that depends on the order. The safe beginner setup is the simple one: store your own provider keys, use them, and ignore credits.

For models that Helicone's registry does not know, routing goes only through your own stored deployments. This is the situation with a custom Azure deployment or a fine-tuned model, and it is another reason to add your provider keys early.

Try it
  1. Run the first-request example with model="gpt-4o-mini", then again with model="gpt-4o-mini/openai".
  2. On the Requests page, compare the provider shown for each request.
  3. Call GET https://ai-gateway.helicone.ai/v1/models with your key and skim the list of names.
both calls succeed and show a provider. The pinned call is guaranteed to use OpenAI, while the bare name may use whichever provider the gateway picked. The models list shows the names you can route to.

Caching: stop paying for the same answer twice

A model answers the same question the same way often enough that repeating the call wastes money. Think of a support bot that gets "what are your opening hours?" a hundred times a day, or of your own test suite that re-runs identical prompts on every commit. Caching stores the first response and returns it for later identical requests, instantly and at no provider cost.

It is off by default. You switch it on with one header:

caching.py
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "What are your opening hours?"}],
    extra_headers={"Helicone-Cache-Enabled": "true"},
)

Send that same request twice. The first is a miss and goes to the provider. The second returns from the cache. To check which one you got, read the response header Helicone-Cache, which says HIT or MISS.

The cache is tunable with a few more headers:

Header Default Meaning
Cache-Control: "max-age=3600" 7 days How long an entry lives, in seconds, up to 365 days
Helicone-Cache-Bucket-Max-Size: "3" 1 Keep up to N different responses and return one at random (maximum 20)
Helicone-Cache-Seed: "user-123" none A namespace, so different seeds never share entries
Helicone-Cache-Ignore-Keys: "request_id,timestamp" none Body fields to leave out when deciding whether two requests match

Understand how a hit is decided, because it explains nearly every caching surprise. Helicone builds a cache key by hashing the seed, the full URL, the full request body, and some headers including the authorisation. Two requests hit the same entry only if all of that is identical. One changed character in the prompt, one different temperature, or one timestamp buried in the body gives a different key, and therefore a miss.

The bucket setting deserves a sentence. With a bucket size of one, the same question always returns the same stored answer, which is what you want for deterministic things. With a larger bucket, Helicone stores several answers to the question and returns one at random, preserving some of the variety you would get from a live model.

Where do cached answers live? On Helicone's edge network, in Cloudflare's key-value storage, not in your own infrastructure. For a team with data-residency rules, such as many Gulf and Egyptian employers with regulated customer data, that is a fact to raise with whoever owns compliance before caching prompts that contain personal data.

A cache that never hits If every call is a miss, compare two requests byte for byte. Usual causes: a timestamp, a request ID or a random value inside the body; a different temperature; a different cache seed; or a header value sent as a real boolean instead of the string "true". Use Helicone-Cache-Ignore-Keys for fields that should not count, and always send strings.
Try it
  1. Send the opening-hours request with caching enabled, twice, printing the Helicone-Cache header each time (with curl -i).
  2. Change one word in the prompt and send it again.
  3. Look at the dashboard: which requests show a cache hit, and what did the hit cost?
first MISS, second HIT at effectively zero cost and much lower latency, and the edited prompt is a MISS again because its cache key differs.

Retries and rate limits

Two more gateway features protect you from failures and from yourself.

Retries

Model providers sometimes fail briefly: a 429 because they are busy, a 503 during a hiccup. Retrying after a short wait usually succeeds. Helicone can do that for you:

retries.py
response = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Write a haiku about the Nile."}],
    extra_headers={
        "Helicone-Retry-Enabled": "true",
        "Helicone-Retry-Num": "3",
        "Helicone-Retry-Factor": "2",
    },
)

Helicone-Retry-Num is how many retries (default 5), Helicone-Retry-Factor is the multiplier for exponential backoff, meaning each wait is that many times longer than the last (default 2), and the minimum and maximum waits default to 1000 and 10000 milliseconds. Retries happen on status 429, 500, 502, 503 and 504, and never on other 4xx errors, which would fail again identically: a malformed request does not fix itself by being repeated. Each attempt is logged separately, so you can see the retries in the dashboard.

The docs add a useful piece of guidance: if you use the gateway's automatic provider routing, failover to another provider is usually better than retrying the same one. Retries earn their place when you have pinned a single provider.

Rate limits

A rate limit caps how much traffic passes in a window of time. The use is protection: a runaway loop, an abusive user, or a bug that would otherwise spend your month's budget in an hour. The header has a compact format:

TEXT
Helicone-Rate-Limit-Policy format
Helicone-RateLimit-Policy: [quota];w=[seconds];u=[request|cents];s=[user|property]

Read it piece by piece. quota is the allowed amount. w is the window length in seconds, and it must be at least 60. u is the unit, either request (the default) or cents for a spending cap. s is the segment, meaning what to count separately, such as per user or per property value.

rate_limits.py
# 1000 requests per hour, across everyone
extra = {"Helicone-RateLimit-Policy": "1000;w=3600"}

# 100 requests per day, per user (needs the user header)
extra = {
    "Helicone-RateLimit-Policy": "100;w=86400;s=user",
    "Helicone-User-Id": "user-123",
}

# 500 cents ($5) per hour, per user
extra = {
    "Helicone-RateLimit-Policy": "500;w=3600;u=cents;s=user",
    "Helicone-User-Id": "user-123",
}

When a call exceeds the policy, Helicone returns 429, and the response headers Helicone-RateLimit-Limit and Helicone-RateLimit-Remaining tell you the state. Your code should treat that 429 like any other: back off, show a friendly message, or queue the work.

A limit that never seems to apply A per-user policy (s=user) does nothing unless the same request also carries Helicone-User-Id. A per-property policy (s=Organization, say) does nothing unless Helicone-Property-Organization is present. Without the matching header there is nothing to count against, so the request is effectively unlimited. Token-based limits and several policies on one request are listed as "coming soon", so do not rely on them.

Both features live in the gateway, so both need proxy integration. If you chose async logging, neither exists for you, and you must implement retries and limits in your own code.

Try it
  1. Set a policy of "2;w=60" on a request, and send it four times quickly.
  2. Note the status codes and read the Helicone-RateLimit-Remaining header on each.
  3. Wait a minute and send once more.
the first two succeed with remaining counting down, the next are 429 with remaining 0, and after the window passes the request succeeds again.

Feedback, scores and finding requests later

Observability is not only watching; it is also judging. Was the answer any good? Helicone lets you attach a verdict to a request after the fact, and then query your history programmatically.

The glue is the request ID. Either read it from the helicone-id response header, or choose your own by sending Helicone-Request-Id with a UUID you generate, which is easier because you know the ID before the call returns.

Feedback is a thumbs-up or thumbs-down from a person, stored as a boolean. Scores are named numbers or booleans from any source, for example an automated check. Both go to the REST API at https://api.helicone.ai (the EU region uses https://eu.api.helicone.ai), authenticated with your Helicone key as a bearer token:

BASH
# A user clicked thumbs-up
curl -X POST "https://api.helicone.ai/v1/request/$REQUEST_ID/feedback" \
  -H "Authorization: Bearer $HELICONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"rating": true}'

# An automated check produced scores
curl -X POST "https://api.helicone.ai/v1/request/$REQUEST_ID/score" \
  -H "Authorization: Bearer $HELICONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"scores": {"accuracy": 92, "is_safe": true}}'

There is a trap in the second call. Scores must be integers or booleans. A float such as 0.92 is rejected. The convention is to multiply by 100 and send 92. Scores are also aggregated with a delay of roughly ten minutes, so do not panic if they do not appear instantly.

You can add a property to an existing request as well, with PUT /v1/request/{requestId}/property and a body like {"Environment": "production"}.

To pull history out, the request-query endpoint accepts a filter:

BASH
curl -X POST "https://api.helicone.ai/v1/request/query-clickhouse" \
  -H "Authorization: Bearer $HELICONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "filter": {
      "request_response_rmt": {
        "properties": { "Environment": { "equals": "production" } }
      }
    },
    "limit": 10,
    "offset": 0
  }'

Study the filter shape, because it is the classic way to get an empty list. The property condition must be wrapped inside request_response_rmt, which is the name of the underlying table. If you write the property filter at the top level, the API does not complain: it returns [], and you will wonder where your data went. Conditions combine with {"left": ..., "operator": "and", "right": ...} when you need more than one.

Try it
  1. Send a request with your own Helicone-Request-Id, using a UUID from uuidgen or Python's uuid4().
  2. Post a feedback rating of true for that ID.
  3. Open the request on the dashboard and find the feedback, then run the query call with a property filter that matches one of your labels.
the request shows your rating, and the query returns your labelled requests. Remove the request_response_rmt wrapper and run it again to see the silent empty result for yourself.

Streaming, privacy and the regions

Three smaller topics that beginners bump into in their first week.

Streaming and cost

Chat applications stream: the model's words arrive a few at a time instead of all at once. The cost of a streamed response depends on the usage figures the provider sends at the end of the stream, and by default some providers do not send them. The result is a cost of zero or a wrong number in the dashboard. The fix is to ask for usage explicitly:

streaming_usage.py
stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "Tell me about Cairo in two sentences."}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

If you cannot change the call, the header helicone-stream-usage: "true" does the same job. Forgetting this and then trusting the streamed cost column is an easy way to under-report your spend.

Keeping bodies out of the logs

Sometimes a prompt contains something you do not want stored: a national ID, a medical note, a password someone typed into a chat box. Two headers tell Helicone not to keep the bodies:

omit.py
extra_headers={
    "Helicone-Omit-Request": "true",
    "Helicone-Omit-Response": "true",
}

Read the fine print. These stop the bodies being stored, but they are still sent to Helicone's backend and held briefly in memory. They reduce what is kept; they do not keep data from ever reaching Helicone. If a regulation forbids your data leaving your network, these headers are not sufficient, and self-hosting or async logging with content tracing disabled is the conversation to have. Redacting sensitive values in your own code before the call is always the strongest protection, and the simplest.

Regions

Helicone Cloud has a US region (us.helicone.ai) and an EU region (eu.helicone.ai), and your dashboard and REST API live in whichever your organisation was created in. The EU region uses https://eu.api.helicone.ai for the REST API, and forgetting this swap makes correct keys return authentication errors. There is no Middle East region, so teams whose customers' data must stay in the Gulf or in Egypt should treat the choice between EU cloud, self-hosting, or a different tool as a compliance decision, not a technical afterthought.

Try it
  1. Run the streaming example with and without stream_options.
  2. Compare the cost shown on the dashboard for the two requests.
  3. Add the two omit headers to a third request and open it in the dashboard.
the stream without usage may show a missing or zero cost while the other shows a real figure, and the omitted request shows its metadata (tokens, cost, latency) but not its prompt and answer bodies.

Running Helicone yourself

If you have no cloud account, this is how you get a real Helicone to play with. The project publishes one Docker image that bundles the dashboard, the API service, the databases and the file storage, so a single command starts everything. If you are new to containers, the Docker guide explains what the commands below do.

BASH
docker pull helicone/helicone-all-in-one:latest

docker run -d \
  --name helicone \
  -p 3000:3000 \
  -p 8585:8585 \
  -p 9080:9080 \
  helicone/helicone-all-in-one:latest

Three ports are published: 3000 is the dashboard, 8585 is the API service (called Jawn), and 9080 is the object storage where request bodies live. Open http://localhost:3000 in a browser. To confirm the API service is alive, run:

BASH
curl http://localhost:8585/healthcheck

There is a wrinkle that catches everyone. The container has no email service, so after you register an account, the account is unverified and you cannot log in. You verify it by hand, directly in the container's Postgres database:

BASH
docker exec -u postgres helicone psql -d helicone_test -c \
  "UPDATE \"user\" SET \"emailVerified\" = true WHERE email = 'your@email.com';"

If the dashboard then says "No organization ID found", your user exists but belongs to no organisation, and you must create one and add yourself to it with the SQL in the official self-host page. The official page lists the exact statements; copy them from there rather than retyping.

The second wrinkle: data vanishes when the container is removed, because the databases live inside it. To keep your data, mount volumes:

BASH
docker run -d --name helicone \
  -p 3000:3000 -p 8585:8585 -p 9080:9080 \
  -v helicone-postgres:/var/lib/postgresql/data \
  -v helicone-clickhouse:/var/lib/clickhouse \
  -v helicone-minio:/data \
  helicone/helicone-all-in-one:latest

Now the honest part about what you can do with this. The all-in-one image runs the dashboard and the API service. It does not include the LLM gateway: in August 2026 the API service stopped proxying model traffic, and requests to its old /v1/gateway/... routes now return 404. The documentation says model traffic should go through "the AI Gateway, which you can deploy next to this image", but the self-hosted gateway page currently returns Not Found, and the standalone Rust gateway project is abandoned. The documented path for self-hosters to get traffic in is therefore logging rather than proxying, and the exact endpoint configuration for pointing the logging SDKs at your own server is not spelled out in the docs. For a beginner this means a local install is excellent for learning the dashboard, the data model and the operations, but do not expect the gateway features (caching, retries, rate limits) to work against it.

Self-hosting needs care beyond a demo The published image dates from August 2025, and it predates several security fixes the company made in 2026. Its default secrets (BETTER_AUTH_SECRET set to change-me-in-production, MinIO credentials minioadmin) are for local use only. Never expose this container to the internet as it stands. Bind it to your own machine, change the secrets if you leave it running, and treat it as a learning environment.

If you ever run it on a remote server, every URL setting must share one origin. The dashboard URL, the API URL and the storage URL must all be the same host, because mixing localhost with a public IP produces an "Invalid origin" error at sign-in. The variables involved are SITE_URL, BETTER_AUTH_URL, NEXT_PUBLIC_APP_URL, NEXT_PUBLIC_HELICONE_JAWN_SERVICE and S3_ENDPOINT, and NEXT_PUBLIC_IS_ON_PREM must be set to true or you meet an infinite redirect loop. Mid-level and Senior cover production self-hosting.

Try it
  1. Start the container with the three volumes and open http://localhost:3000.
  2. Run the healthcheck curl and confirm it answers.
  3. Register a user, verify the email with the SQL statement, log in, and find the Requests page (it will be empty).
  4. Stop and remove the container, start it again with the same volumes, and confirm your account is still there.
the dashboard loads, the healthcheck answers, and your account survives the restart because the data lives in named volumes. Without the volumes it would be gone.

Configuration and the errors you will actually meet

By now you have seen most of Helicone's configuration. It is almost all headers, plus two secrets. It helps to hold the whole surface in one table.

Want Set Where
Authenticate to the gateway Authorization: Bearer <Helicone key> Client API key
Label a user Helicone-User-Id Header
Label anything else Helicone-Property-<Name> Header
Group a run Helicone-Session-Id, -Path, -Name Headers
Cache Helicone-Cache-Enabled, Cache-Control Headers
Retry Helicone-Retry-Enabled, -Num, -Factor Headers
Limit Helicone-RateLimit-Policy Header
Skip storing bodies Helicone-Omit-Request, Helicone-Omit-Response Headers
Your provider keys Providers page Dashboard

Reading the errors

When something fails, work in this order: read the status code, then read the error message body, then check the dashboard row if one exists. A request that reached Helicone and failed is still logged, which is itself useful evidence.

401 "Authentication failed", type invalid_api_key. The gateway tried to call the provider and the provider key stored in Helicone was rejected. Open the Providers page and re-enter the key. If you are on a legacy URL, the cause is usually the key mix-up: provider key in Authorization, Helicone key in Helicone-Auth: Bearer sk-helicone-..., with the Bearer prefix.

429 Insufficient credits. You have no provider key stored, or no credits, so the gateway could not find a way to pay for the call. Add a provider key.

429 with Helicone-RateLimit-Remaining: 0. This is your own policy working. Wait for the window or change the policy.

403 "Wallet suspended" or a blocked model. An account-level block. Email support; nothing in your code causes it.

400 PROMPT_FLAGGED_FOR_MODERATION or PROMPT_THREAT_DETECTED. You or an admin enabled the optional moderation or security screening headers and a prompt was blocked. Handle these errors in your app as an expected outcome, not a crash.

The cache always says MISS. Compare two requests for any changing field, and send header values as strings.

Retry or other headers seem ignored. A value was sent as a number or boolean. Make it a string.

Query API returns []. The property filter is not wrapped in request_response_rmt.

A score is rejected. It was a float. Send an integer or a boolean.

Streaming cost is zero or off. No usage chunk. Add stream_options={"include_usage": True}.

Cost missing for a custom model. Helicone cannot resolve the model name to a price. The Helicone-Model-Override header supplies a name for cost calculation.

Self-host: the browser shows "connection refused" for API calls. The page is calling localhost:8585 from a machine where that is wrong. Set NEXT_PUBLIC_HELICONE_JAWN_SERVICE to the address browsers can reach.

Self-host: "Invalid origin" or an endless redirect. The origin variables disagree, or NEXT_PUBLIC_IS_ON_PREM=true is missing.

Self-host: "/v1/gateway" returns 404. That route was removed in August 2026. It is not misconfigured.

Sign-up page says "Sign ups are disabled". That is maintenance mode. Self-host instead.

Debug from the outside in When a call misbehaves, reproduce it with curl -i against the gateway before touching your application code. If curl works, the bug is in how your program builds the request. If curl fails the same way, the bug is in the keys, the headers or the provider, and you have saved an hour of reading your own code.
Try it
  1. Deliberately break a working call: change one character of your key.
  2. Read the status and body, then find the request (if any) on the dashboard.
  3. Fix it, then break it a different way: send a cache header with a boolean True instead of the string.
the first break produces a 401 with a readable message, the second produces no error at all, only a cache that never engages. The second is the more dangerous kind of failure, because nothing tells you it happened.

Putting it all together

Let us build one small project that uses everything above: a command-line assistant that answers questions for a fictional customer-support desk, with labelled traffic, a cache, retries, a rate limit, a session, feedback and a spending check. Save it as desk.py. It assumes HELICONE_API_KEY is set and a provider key is stored in Helicone.

desk.py
import os
import sys
import uuid
import requests
from openai import OpenAI

HELICONE_KEY = os.environ["HELICONE_API_KEY"]

client = OpenAI(
    base_url="https://ai-gateway.helicone.ai",
    api_key=HELICONE_KEY,
    default_headers={
        "Helicone-Property-Environment": "local",
        "Helicone-Property-App": "support-desk",
        "Helicone-Cache-Enabled": "true",
        "Helicone-Retry-Enabled": "true",
        "Helicone-Retry-Num": "3",
    },
)

session_id = str(uuid.uuid4())


def ask(user_id: str, question: str) -> tuple[str, str]:
    request_id = str(uuid.uuid4())
    response = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[
            {"role": "system", "content": "You are a concise support agent."},
            {"role": "user", "content": question},
        ],
        extra_headers={
            "Helicone-Request-Id": request_id,
            "Helicone-User-Id": user_id,
            "Helicone-Session-Id": session_id,
            "Helicone-Session-Path": "/answer",
            "Helicone-Session-Name": "Support Desk",
            "Helicone-RateLimit-Policy": "20;w=60;s=user",
        },
    )
    return request_id, response.choices[0].message.content


def rate(request_id: str, good: bool) -> None:
    requests.post(
        f"https://api.helicone.ai/v1/request/{request_id}/feedback",
        headers={"Authorization": f"Bearer {HELICONE_KEY}"},
        json={"rating": good},
        timeout=10,
    )


if __name__ == "__main__":
    user = sys.argv[1] if len(sys.argv) > 1 else "demo-user"
    while True:
        question = input("Question (blank to quit): ").strip()
        if not question:
            break
        rid, answer = ask(user, question)
        print(answer)
        verdict = input("Helpful? [y/n]: ").strip().lower()
        rate(rid, verdict == "y")

Run it, ask the same question twice, rate one answer up and one down, and ask five different questions. Then open the dashboard and inspect the results in this order.

  1. Requests page, filtered by the property App = support-desk. You should see all your calls and only yours.
  2. The repeated question. The second copy should be a cache hit with near-zero cost and latency.
  3. Sessions page, where "Support Desk" shows one session containing every call, with a total cost.
  4. One request, opened. It shows the user ID, the properties, and your feedback rating.
  5. Filter by user to see the spending of demo-user.

Then answer the questions the project was built to answer: which question cost the most, what share of calls were cache hits, and what was the session's total cost. If you can answer those three from the dashboard in under a minute, you have achieved what a bare SDK call could not.

If you are on the local self-hosted install without a gateway, read this project as a design exercise: the labelled headers and the feedback call are the same ideas, but the cache, retry and rate-limit headers need a gateway to act on them.

Try it
  1. Run desk.py and ask six questions, including one repeated.
  2. Answer the three questions above from the dashboard only.
  3. Extend the script: add a Helicone-Property-Feature header with a value per question type, and filter by it.
you can point to the cache hit, the session total and a per-user figure without opening your code. The extension shows how one more header gives you a new axis to slice.

What you can now do, and what comes next

You can now explain what Helicone is and where it stands: an Apache-2.0 LLM observability tool with a gateway, in maintenance mode since March 2026, with sign-ups closed and self-hosting open. You can distinguish the two integration styles and say what each costs you. You can send a request through the AI Gateway with the OpenAI SDK, find it on the dashboard, and read its tokens, cost and latency. You can label traffic with users, properties and sessions, switch on caching, retries and rate limits with string-valued headers, attach feedback and scores to a request, query history with the correct filter shape, cope with streaming cost, and keep bodies out of storage with a clear understanding of the limits of that. You can run the dashboard locally and explain honestly what a local install does and does not give you. And you can read the common errors by their status code.

That is a solid foundation, and it is also the set of things people get wrong in interviews. The interview guide for this level turns it into questions and model answers, and the tips guide collects the mistakes that cost beginners the most time.

What comes next, at Mid-level: prompt management with versioned templates and environments, provider routing and fallback chains in depth, async logging for custom models, alerts and webhooks, the query language, and how to wire Helicone into an agent framework. At Senior: architecture and failure modes, security and the trust model, data export and migration plans, multi-tenant attribution, upgrades when there are no maintained releases, and how to argue for or against keeping a tool in maintenance mode.

Two neighbouring guides in this catalogue are worth reading alongside. Langfuse is an observability platform with a similar purpose and is a frequent migration target, and Docker is the foundation for the self-hosting steps. Reading them next to this one shows you which ideas are Helicone's and which are simply how LLM observability works.

Try it
  1. Without looking back, write the eight headers you would add to a production call: two labels, one session set, one cache, one retry, one limit, one omit.
  2. Write one sentence on the risk of depending on a tool in maintenance mode, and one mitigation.
a list of exact header names with string values, and a risk and mitigation such as "no new features; export data regularly and keep the integration behind one base-URL setting so it can be swapped".

Sources