Skip to content
Back to student guides
OpenTelemetryDevOpsObservability3 levels105 sectionsCovers OpenTelemetry Collector 0.162

The Complete OpenTelemetry Guide

Instrument services with vendor-neutral traces, metrics and logs using OpenTelemetry. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
18sections
35examples

This is part one of three. It covers everything you need to start doing real work with OpenTelemetry, not a teaser. By the end you can explain what OpenTelemetry is and is not, run a Collector on your laptop, read a config file for it without fear, add tracing to a small Python service without touching its code, add your own spans and metrics by hand, and diagnose the handful of mistakes that stop nearly every beginner's first attempt. Mid-level and Senior take the same topics further; nothing here is thrown away.

A note on versions before we begin. OpenTelemetry is not one program with one version number. It is a specification, a wire protocol, a set of conventions, a program called the Collector, and a separate software development kit (SDK) for every programming language, each released on its own schedule. When this guide gives a version it means the Collector v0.162.0, released at the end of September 2026, alongside version 1.61.0 of the specification and version 1.45.0 of the Python SDK. The Collector is still a 0.x program, which means any release may rename or remove a setting. That is why you will see a few notes below of the form "older tutorials show X, current releases want Y". Trust the current form.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and telemetry only becomes real when you watch your own request turn into a tree of timed spans on your own screen.

What OpenTelemetry is, and the problem it solves

When a program runs on your laptop you understand it by reading the code and adding print statements. When it runs on a server you cannot attach to, serving thousands of users, calling a database, a cache, a payment provider and three other services, print statements stop working. You need the running system to tell you what it is doing. That ability is called observability: the degree to which you can work out what is happening inside a system from the data it emits.

The data comes in a few shapes that we will meet one by one: traces (the story of one request), metrics (numbers measured over time), and logs (timestamped text records). Collectively these are called telemetry. OpenTelemetry, usually shortened to OTel, is the open standard and toolkit for generating, collecting and exporting that telemetry. It belongs to the Cloud Native Computing Foundation, the same foundation that hosts Kubernetes and Prometheus.

The most important sentence in this guide is the next one. OpenTelemetry is not a backend. It does not store your data and it does not draw you a dashboard. It produces telemetry in a standard shape and delivers it to whatever you choose to store and visualise it: Jaeger, Grafana Tempo, Prometheus, Loki, Elasticsearch, or a commercial vendor. Beginners who install OpenTelemetry and then ask "where do I look at my traces?" have missed this. You look at them in a backend you pick separately.

YOUR CODEinstrumented
→
SDKin the process
→
COLLECTORreceive, process, export
→
BACKENDJaeger, Tempo, vendor

The diagram is the whole architecture. Everything in the next eight thousand words is a detail of one of those four boxes.

What came before

Before OpenTelemetry, every monitoring product shipped its own library. If you wanted Datadog you installed the Datadog agent and its tracing library; if you wanted New Relic you ripped that out and installed theirs; if you preferred open source you used OpenTracing for traces and OpenCensus for metrics, two competing projects that solved overlapping halves of the problem. The consequence was lock-in. Your application code was littered with calls to one vendor's API, so changing vendor meant re-instrumenting every service. Worse, each library invented its own names: one called the HTTP status code http.status, another statusCode, another response_code, which made it impossible to compare data between tools.

OpenTelemetry was formed by merging OpenTracing and OpenCensus. It answers the lock-in problem with two separations. First, it separates how you produce telemetry from where you send it: you instrument your code once against the OpenTelemetry API, and you change the destination by changing configuration, not code. Second, it standardises the names of things. A set of rules called semantic conventions says that an HTTP status code is always http.response.status_code, a database system is always db.system.name, and so on, so data from a Python service and a Java service line up.

The project has a few consequences you should notice now.

Instrumentation and backend are decoupled. You can send the same data to two backends at once, or move from a self-hosted tool to a vendor and back, without a code change. For a team in the Gulf or in Egypt weighing data-residency rules, this matters in a practical way: you can keep telemetry inside a regional cloud or on your own servers today, and change your mind later without touching application code.

There is a lot of it, and it moves quickly. The Collector releases roughly every two weeks. Tutorials from a year ago may show component names that are now deprecated. This guide uses the current names and tells you when an old one is still accepted.

Not everything is equally mature. Traces and metrics are stable in the major languages. Logs are stable in some languages and still in development in others, including Python and JavaScript at the time of writing. Profiles, a fourth signal, are alpha. We will be honest about which parts are solid as we go.

Try it
  1. Think of a service you have used or built that calls at least two other things, such as a database and an external API.
  2. Write down three questions you could not answer if it ran slowly in production: for example "which of the calls was slow for this one user?"
  3. Mark each question as needing one request's story or a number over time.
a mix of both kinds. The first kind is what traces answer, the second is what metrics answer, and the reason OpenTelemetry carries several signals rather than one.

Traces, metrics and logs: the three signals

A signal is a category of telemetry. OpenTelemetry has three you will use daily, plus a newer fourth. Understanding what each is good at tells you which to reach for.

Traces

A trace is the path of a single request through your system. Suppose a user taps "Pay" in a mobile app. The app calls an API gateway, which calls an orders service, which calls an inventory service and a payments service, which each talk to a database. A trace records that whole journey, and because it follows one request it can tell you something no average can: this request was slow, and it was slow here.

A trace is made of spans. A span is one timed operation: "handle POST /pay", "SELECT from orders", "call the payment provider". Every span in a trace shares the same 16-byte trace ID. Each span also has its own 8-byte span ID and, except for the first, the ID of its parent. Follow the parent links and you get a tree. The span with no parent is the root span, usually the first service to receive the request.

TEXT
POST /pay                                  [=================================] 420 ms
  validate cart                            [==]                                 30 ms
  inventory.reserve                           [=======]                         90 ms
    SELECT stock                                [===]                           40 ms
  payments.charge                                    [========================] 260 ms
    POST https://psp.example/charge                    [====================]   230 ms

Read the picture the way you would read a waterfall in your browser's network tab. Time runs left to right, indentation shows who called whom, and the widest bar near the bottom is the place to look. Here the payment provider took 230 of the 420 milliseconds. Nobody had to guess.

What does a span carry? A name, start and end timestamps, a kind (SERVER when it handles an incoming request, CLIENT when it makes an outgoing one, PRODUCER and CONSUMER for messaging, INTERNAL for work inside one process), a status (Unset, Ok or Error), and three kinds of extra detail. Attributes are key-value pairs that describe the operation, such as http.response.status_code = 200 or order.id = "123". Events are timestamped notes inside the span, such as "cache miss" or an exception with its stack trace. Links point to spans in other traces, for cases like a batch job that processes messages from many different requests.

Metrics

A metric is a measurement aggregated over time: requests per second, the 95th-percentile latency, the number of items in a queue, memory in use. Where a trace answers "what happened to this request?", a metric answers "how is the system doing overall?". Metrics are cheap, because a million requests still become a handful of numbers per minute, which is why alerts and dashboards are built on them.

You create metrics through instruments, and choosing the right instrument is the one piece of metrics vocabulary worth learning now. A counter only goes up: orders processed, bytes sent. An up-down counter can go up and down: items currently in a queue. A histogram records a distribution of values, such as request durations, so you can later ask for percentiles. A gauge records the latest value of something, such as a temperature. Counters, up-down counters and gauges also have asynchronous forms, where instead of calling the instrument when something happens you register a callback that the SDK invokes when it collects data, which suits things like "current memory use" that you read on demand.

Logs as the third signal

A log record is a timestamped message, usually with a severity such as INFO or ERROR, a body, and attributes. Logs are the oldest signal and every application already produces them. OpenTelemetry's contribution is not a new way to write logs. It connects the logging you already have to the rest of your telemetry. Through a bridge or appender attached to your existing logging library (Python's logging, Java's Logback, and so on), log records are emitted in the OpenTelemetry format, and, crucially, each one is stamped with the trace ID and span ID of the request that was running when the line was written. Click from a trace to the log lines from that exact request, or the reverse. That correlation is the payoff.

The fourth signal and the odd one out

Profiles, continuous records of where a program spends CPU time, are the newest signal. They are in alpha, so this guide mentions them and moves on. Baggage is sometimes mistaken for a signal, but it is a mechanism: key-value pairs that travel along with a request from service to service so that, say, a customer tier set at the front door is visible deep in the call chain. Baggage is not added to your spans automatically. You have to copy it into attributes yourself, a trap we return to later.

Why not just logs? You can reconstruct a request from logs if every service logs a shared request ID and you are willing to search across all of them by hand. Traces do that joining for you, with timing. Metrics do the aggregation that logs make expensive. The three signals are not rivals. They are three views onto the same system, and OpenTelemetry's job is to keep them connected through shared IDs and shared resource attributes.
Try it
  1. Take the "Pay" example above and, without tools, sketch the trace tree for a request to your own project's most important endpoint.
  2. For each box, guess the duration and mark the widest one.
  3. Then name one metric you would want on that endpoint and which instrument (counter, up-down counter, histogram, gauge) you would use.
a latency metric comes out as a histogram, a request total as a counter, and a "currently in flight" number as an up-down counter. If your guesses put the widest bar in the wrong place, that is exactly why measuring beats guessing.

The mental model: resource, scope, context

Three more nouns give you most of the model. Learn them properly now and the rest of the guide is easy to read.

Resource: who is speaking

A resource is a set of attributes describing the thing that produced the telemetry: the service, the host, the container, the Kubernetes pod, the cloud region. Every span, metric point and log record from one process carries the same resource. The most important attribute in all of OpenTelemetry is service.name. It is how your backend groups data: without it, everything arrives labelled unknown_service (or unknown_service:java, unknown_service:python), and the first thing you notice in a dashboard is a service with that useless name. Other commonly set attributes are service.version, service.namespace, service.instance.id and deployment.environment.name (the current name; the older deployment.environment is replaced).

You set them with two environment variables that every language honours:

BASH
export OTEL_SERVICE_NAME=checkout
export OTEL_RESOURCE_ATTRIBUTES=service.version=1.4.2,deployment.environment.name=staging

The second variable takes comma-separated key=value pairs. Resource detectors can also fill in host, container, process and cloud attributes automatically, which is why you rarely type them by hand in production.

Instrumentation scope: which code spoke

The instrumentation scope is the name, and optionally the version, of the library or module that created the telemetry. When your code asks for a tracer you pass a name such as checkout.orders, and every span made through that tracer carries it. This lets you tell a span created by the Flask instrumentation library from one created by your own business code, even though both belong to the same service.

Context and propagation: how a trace survives the network

Inside one process, the context is an immutable carrier that knows which span is currently active. When you start a span inside another, the SDK uses the context to set the parent automatically. You do not pass anything around by hand.

Across a network boundary the context cannot travel by itself, so it is written into the request and read back out on the other side. This is context propagation, performed by a propagator that injects context into a carrier (the HTTP headers of an outgoing request) and extracts it on the receiving end. The default standard is W3C Trace Context, which uses a header called traceparent:

TEXT
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

The four dash-separated fields are the version (00), the 32-hex-digit trace ID, the 16-hex-digit ID of the calling span, and two hex digits of flags, where the last bit says "this trace is being sampled". When service B receives that header it creates its spans with the same trace ID and the caller's span as parent. That is the entire mechanism that stitches many services into one trace. If any service in the chain drops the header, the trace splits into two unrelated traces, which is the most common cause of "my trace is broken" questions.

A broken trace is usually a dropped header If you see two short traces where you expected one long one, look at the hop between them. A proxy that strips unknown headers, a message queue where nobody copied the context into the message, or a service that was never instrumented will all cut the thread. Propagation fails silently. Nothing reports an error.
Try it
  1. Split the example traceparent above into its four fields by hand.
  2. Which field would be identical in every service that handles the request, and which would be different?
  3. Set OTEL_SERVICE_NAME in your shell to the name of a project of yours.
the trace ID is the same everywhere; the span ID changes at each hop. You have just set the most important resource attribute in OpenTelemetry.

The pieces: API, SDK, OTLP, Collector

The phrase "OpenTelemetry" covers several distinct parts, and confusing them is the next most common beginner stumble. Here they are in the order data flows.

The API is the set of calls your code makes: "start a span", "add one to this counter", "emit this log record". The API is deliberately a no-op until something installs an implementation: without an SDK, the calls do nothing and cost almost nothing. That design is what lets a library author, say the maintainer of an HTTP client, add OpenTelemetry calls without forcing any overhead or dependency on users who do not care. The rule is that libraries depend only on the API.

The SDK is the implementation of the API that your application installs and configures. It decides what happens to a span once it ends: whether it is sampled at all, how it is batched, and which exporter sends it. The SDK is entered through providers (TracerProvider, MeterProvider, LoggerProvider), normally registered globally once at startup.

An exporter serialises telemetry and sends it somewhere. The standard destination is another process speaking OTLP.

OTLP, the OpenTelemetry Protocol, is the wire format: protobuf messages carried over gRPC on port 4317 or over HTTP on port 4318. The HTTP flavour uses the paths /v1/traces, /v1/metrics and /v1/logs. Commit this pair to memory: 4317 is gRPC, 4318 is HTTP. Sending one protocol to the other's port is the number-one first-day bug, and we will meet it again in the errors section.

The Collector is a standalone program that receives telemetry, processes it and exports it onward. It is optional, since an SDK can send straight to a backend, but almost every real setup includes one, for reasons we unpack in the next section. Semantic conventions are not software at all: they are the agreed attribute and metric names that keep data comparable.

And then there is instrumentation, the act of adding telemetry. It comes in three forms worth distinguishing:

  • Code-based (manual) instrumentation: you call the API yourself, to create a span around your own business logic.
  • Instrumentation libraries: packages that wrap popular frameworks (Flask, Express, a JDBC driver) and emit standard spans for you.
  • Zero-code (automatic) instrumentation: an agent that attaches to your program at startup and applies the instrumentation libraries without you editing source. Python has opentelemetry-instrument, Java has a -javaagent JAR, and Node.js has a register hook.

Start with zero-code, add manual spans where the automatic ones do not describe your business, and you will have good coverage with little work.

Which language? The examples below use Python because it needs the least ceremony, and because most readers of this guide work in it. Everything about the Collector, OTLP and the environment variables is identical for Java, Node.js, Go and .NET. Only the SDK calls differ.
Try it
  1. Draw the four boxes from the first diagram on paper.
  2. Label each of these on the correct box: API, SDK, exporter, OTLP, Jaeger, batching, sampling, semantic conventions.
  3. Note which two of them would stay the same if you switched from Jaeger to a commercial vendor.
your instrumented code (API plus semantic conventions) stays untouched; only the final box and the exporter destination change. That is the decoupling OpenTelemetry sells.

Why a Collector, and where it sits

If an SDK can send OTLP directly to a backend, why run an extra program? Four reasons, all practical.

First, offloading. The SDK inside your application should finish quickly and return to work. Handing data to a nearby Collector on the same machine is faster than negotiating with a distant backend, and the Collector absorbs the retries, backoff and buffering so that your application does not.

Second, one place for credentials and policy. If every service holds a vendor API key, rotating it means redeploying everything. With a Collector, only the Collector holds the key. The same goes for scrubbing personal data, dropping noisy health-check spans, and adding Kubernetes labels: do it once, centrally.

Third, fan-out and format conversion. One Collector can receive OTLP and send traces to Jaeger, metrics to Prometheus and everything to a vendor at the same time. It can also receive other formats, such as Prometheus scrape targets, Zipkin or Jaeger spans, and convert them.

Fourth, changing your mind later costs a config edit in one place, not a change in every application.

There are three common shapes. With no Collector, the SDK exports directly to the backend: simplest for a demo, weakest for production. As an agent, a Collector runs next to each application, as a sidecar, a Kubernetes DaemonSet or a service on the host. As a gateway, a central pool of Collectors sits behind a load balancer. Many production setups use both, agents close to the applications and a gateway in front of the backends. As a beginner you run one Collector on your laptop and that is precisely the right size.

The Collector's own model

Inside the Collector there are five kinds of component, and once you know them you can read any config in the world.

  • A receiver gets data in. Some receivers listen for pushes (the otlp receiver), and some pull or scrape (the prometheus and host_metrics receivers).
  • A processor changes data on its way through: batching it, limiting memory, dropping things, adding attributes. Order matters.
  • An exporter sends data out.
  • A connector is an exporter of one pipeline and a receiver of another, and can change the signal type, for example turning traces into metrics. You will not need one yet.
  • An extension provides capabilities that are not part of the data path, such as a health-check endpoint.

Components are wired together in pipelines. There is one pipeline type per signal (traces, metrics, logs), and each pipeline is a list of receivers, then processors, then exporters.

RECEIVERSotlp, prometheus
→
PROCESSORSmemory_limiter, batch
→
EXPORTERSdebug, otlp_grpc

Two rules explain more confusing behaviour than anything else. First, defining a component is not enabling it: a receiver, processor or exporter that appears at the top of the file but is not listed under service.pipelines simply does nothing. Second, each component has an ID of the form type or type/name, for example otlp_grpc/tempo, so that you can have two exporters of the same type pointing at different places.

Try it
  1. Imagine you must send traces to Jaeger and metrics to Prometheus from one Collector.
  2. Write, in words, the two pipelines: which receiver, which processors, which exporter each uses.
two pipelines that share the same otlp receiver but use different exporters. Receivers and exporters may be shared between pipelines; that shape is the whole point of the design.

Installing the Collector and checking it works

There are two official builds you need to know about. The core distribution (otelcol) contains a small set of components. The contrib distribution (otelcol-contrib) contains nearly everything the community maintains, including the file_log, host_metrics, prometheus and k8s_attributes components. For learning, use contrib, so that a component you find in a tutorial is actually present. For production you will eventually build a slimmer custom distribution, which is a Mid-level and Senior topic.

All artifacts live on the collector-releases GitHub page for the version. Replace the version below if a newer one exists, and always pin an exact version; never use latest, because the Collector is 0.x and releases can break configs.

With Docker (any operating system)

This is the quickest route, and the one used for the rest of the guide. Create a file called config.yaml (we write its content in the next section), then run:

BASH
docker run --rm \
  -p 127.0.0.1:4317:4317 -p 127.0.0.1:4318:4318 -p 127.0.0.1:13133:13133 \
  -v "$(pwd)/config.yaml:/etc/otelcol-contrib/config.yaml" \
  otel/opentelemetry-collector-contrib:0.162.0

The -p flags publish the OTLP ports to your machine; the -v flag mounts your file over the image's default config path. The image runs the Collector as its main process, so the terminal is held by its log output. If Docker itself is new to you, the Docker guide covers the flags above.

On Linux, macOS or Windows without Docker

On Debian or Ubuntu, download the .deb package and install it, which also registers a systemd service:

BASH
wget https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/v0.162.0/otelcol-contrib_0.162.0_linux_amd64.deb
sudo dpkg -i otelcol-contrib_0.162.0_linux_amd64.deb

The service reads /etc/otelcol-contrib/config.yaml. After you edit that file, run sudo systemctl restart otelcol-contrib, because the Collector does not reload its config from disk on its own, and read its logs with sudo journalctl -u otelcol-contrib. On Red Hat family systems use the .rpm. On macOS there is no official Homebrew formula for the Collector, so download the tarball for your chip and run it directly:

BASH
curl --proto '=https' --tlsv1.2 -fOL https://github.com/open-telemetry/opentelemetry-collector-releases/releases/download/v0.162.0/otelcol-contrib_0.162.0_darwin_arm64.tar.gz
tar -xvf otelcol-contrib_0.162.0_darwin_arm64.tar.gz
./otelcol-contrib --config=config.yaml

If macOS refuses to run the unsigned binary, remove the quarantine flag with xattr -d com.apple.quarantine ./otelcol-contrib. On Windows there are MSI installers that create a Windows service, plus a plain .tar.gz you can unpack and run as .\otelcol-contrib.exe --config=config.yaml.

Checking the installation

The Collector gives you four quick checks before you ever send it data. Use the binary name that matches what you installed.

BASH
otelcol-contrib --version
otelcol-contrib components
otelcol-contrib validate --config=config.yaml

--version confirms which build you have. components lists every receiver, processor, exporter, connector and extension compiled into this binary, along with its stability level; it is how you answer "is file_log in this build?" in five seconds. validate parses your config and checks it, exiting non-zero with the reason if it is wrong; run it in the habit of a compiler run, before you start the program. With Docker you can run the same commands by appending them to the image name, for example docker run --rm otel/opentelemetry-collector-contrib:0.162.0 components.

When you finally start the Collector with a valid config, look for this line near the end of the startup log:

TEXT
Everything is ready. Begin running and processing data.

If you see it, the receivers are listening. On shutdown (Ctrl+C) you will see Shutdown complete.

Localhost is the default, and Docker changes the meaning of it Since recent releases every Collector server endpoint defaults to localhost rather than 0.0.0.0. Inside a container, "localhost" means the container itself, so a receiver bound to it is unreachable from your machine even though the -p mapping exists. You get "connection refused" with no error in the Collector's log. In Docker and Kubernetes always set the receiver endpoints to 0.0.0.0:4317 and 0.0.0.0:4318, as the config below does.
Try it
  1. Pick an install route and run --version, then components.
  2. Find otlp in the receivers list and debug in the exporters list, and read the stability shown next to them.
  3. Search the output for file_log. Would it be present in the core build?
OTLP and debug are present in both distributions, while file_log only appears in contrib. That difference explains the classic "unknown type" error you will see in the errors section.

Your first Collector configuration

The Collector is configured by one YAML file. The structure is a mirror of the model you just learned: a top-level section for each kind of component, then a service section that switches some of them on.

Here is a complete, working first config. Save it as config.yaml.

config.yaml
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
    spike_limit_percentage: 25
  batch: {}

exporters:
  debug:
    verbosity: detailed

extensions:
  health_check:
    endpoint: 0.0.0.0:13133

service:
  extensions: [health_check]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [debug]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [debug]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [debug]

Read it top to bottom, because every line answers a question.

The otlp receiver opens two listeners, one for gRPC on 4317 and one for HTTP on 4318, bound to 0.0.0.0 so that they are reachable from outside a container. A single receiver handles all three signals on both protocols, which is why the three pipelines below can share it.

The memory_limiter processor is a safety valve. It checks memory use every second, and when the Collector climbs past a soft limit (80 percent of available memory minus the 25 percent spike allowance) it starts refusing new data rather than being killed by the operating system. Two beginner traps live here. Its check_interval has no useful default, so you must set it, and you must set either limit_mib or limit_percentage; leaving them out produces startup errors such as 'check_interval' must be greater than zero. Place it first in every pipeline so that it can protect everything after it.

The batch processor collects data into groups before export, which is far cheaper than sending each span alone. batch: {} means "use the defaults": a batch is sent when it reaches 8192 items or after 200 milliseconds, whichever comes first. Put it last among the processors. One forward-looking note: the project is moving batching into the exporters themselves (a sending_queue with a batch block), and a newer processor called queue_batch exists in development. The batch processor remains supported in this release and is the right thing to learn first.

The debug exporter prints telemetry to the Collector's own log. It is how you see that data arrived. The verbosity levels are basic, normal and detailed, and detailed dumps every attribute. You may find old tutorials using an exporter called logging. That was removed, and using it stops the Collector with the message the logging exporter has been deprecated, use the debug exporter instead.

The health_check extension serves a small HTTP endpoint you can curl, or point a Kubernetes probe at, to ask whether the Collector is up. It is an extension because it is not part of any data pipeline.

Finally service turns things on. The extensions list enables health_check. Each pipeline lists receivers, then processors in execution order, then exporters. If you deleted the service block, nothing above would run.

Defined is not enabled The single most common config mistake is adding a processor or exporter at the top of the file and forgetting to reference it under service.pipelines. The Collector does not complain about an unused definition. It just does not use it. Conversely, referencing an ID that was never defined is an error: service::pipelines::traces: references exporter "otlp_grpc/tempo" which is not configured.

Start the Collector with this config, then, from another terminal, check it is alive:

BASH
curl http://localhost:13133/

A small JSON status document comes back. You now have an idle, healthy Collector. The next step is to give it something to receive.

Sending fake telemetry

The project ships a tool called telemetrygen that generates synthetic data. It is a Go program, so you need Go installed to fetch it:

BASH
go install github.com/open-telemetry/opentelemetry-collector-contrib/cmd/telemetrygen@latest
telemetrygen traces --otlp-insecure --traces 3

--otlp-insecure tells it to use plaintext gRPC, matching the plain 4317 listener. In the Collector's terminal you will see the debug exporter print three traces in detail, each with spans, attributes such as network.peer.address, and the resource with the service name telemetrygen. The first time you see a wall of output appear because of a command you typed, the whole idea clicks: a producer sent OTLP to a receiver, a pipeline handled it, an exporter emitted it.

Try it
  1. Save the config above and run validate on it, then start the Collector.
  2. Curl the health endpoint, then run telemetrygen and find one span's name and attributes in the output.
  3. Break the config on purpose: change the pipeline's exporter to debugg and run validate again.
the typo produces an error naming the pipeline and the unknown exporter ID. Learning to read that message now is worth more than any amount of reading about it.

Your first instrumented app, with no code changes

Now the producing side. We will instrument a tiny Python web service using zero-code instrumentation, which means we do not edit the service at all.

Create a folder with a virtual environment and a file called app.py:

app.py
from flask import Flask

app = Flask(__name__)


@app.route("/hello/<name>")
def hello(name):
    return {"greeting": f"Hello, {name}"}


@app.route("/slow")
def slow():
    import time
    time.sleep(0.4)
    return {"status": "finally"}


if __name__ == "__main__":
    app.run(port=8080)

Install Flask and the OpenTelemetry tooling:

BASH
pip install flask opentelemetry-distro opentelemetry-exporter-otlp
opentelemetry-bootstrap -a install

The first package, opentelemetry-distro, is a convenience bundle: it brings the SDK and the opentelemetry-instrument launcher and sets sensible defaults. The second, opentelemetry-exporter-otlp, is the exporter that speaks OTLP. The third command, opentelemetry-bootstrap, is the clever one. It inspects the packages installed in your environment and installs the matching instrumentation library for each: it sees Flask and installs the Flask instrumentation. If you use uv instead of pip, plain -a install misbehaves; use uv run opentelemetry-bootstrap -a requirements | uv add --requirement -.

Now run the service through the launcher, with the environment variables that say who it is and where to send data:

BASH
export OTEL_SERVICE_NAME=hello-api
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
opentelemetry-instrument python app.py

Nothing about app.py changed. opentelemetry-instrument starts Python, loads the SDK and instrumentation before your code runs, and wraps Flask so that every incoming request creates a span. Call it:

BASH
curl http://localhost:8080/hello/aya
curl http://localhost:8080/slow

In the Collector's terminal, after a few seconds (the SDK batches and sends on a timer), spans appear. A trimmed version of what the detailed output looks like:

TEXT
Resource attributes:
     -> service.name: Str(hello-api)
     -> telemetry.sdk.language: Str(python)
     -> telemetry.sdk.name: Str(opentelemetry)
ScopeSpans #0
InstrumentationScope opentelemetry.instrumentation.flask 0.66b0
Span #0
    Trace ID       : 4bf92f3577b34da6a3ce929d0e0e4736
    Parent ID      :
    ID             : 00f067aa0ba902b7
    Name           : GET /slow
    Kind           : Server
    Attributes:
         -> http.request.method: Str(GET)
         -> http.response.status_code: Int(200)
         -> url.path: Str(/slow)

Read it with the vocabulary you now own. The resource shows service.name: hello-api, which came from the environment variable. The instrumentation scope names the library that created the span, here the Flask instrumentation. The span has a name (GET /slow), a kind of Server, and attributes using standard semantic-convention keys. The parent ID is empty, so this is the root span of its trace. The exact IDs and version strings will differ on your machine. The structure will not.

The protocol must match the port The Python distro defaults to gRPC (OTEL_EXPORTER_OTLP_PROTOCOL=grpc), so port 4317 is correct above. The language-neutral specification default is http/protobuf, and the Java agent and JavaScript SDKs follow the spec, so they want port 4318. If your app reports UNAVAILABLE or the Collector shows nothing, check that the protocol and the port agree before you check anything else.

What about Java and Node.js?

The pattern is the same; only the launcher differs. For Java you download the agent JAR once and attach it with a flag, leaving your application untouched:

BASH
java -javaagent:./opentelemetry-javaagent.jar -Dotel.service.name=hello-api -jar app.jar

For Node.js you install @opentelemetry/api and @opentelemetry/auto-instrumentations-node, then preload the register hook so it loads before your libraries:

BASH
npm install --save @opentelemetry/api @opentelemetry/auto-instrumentations-node
export NODE_OPTIONS="--require @opentelemetry/auto-instrumentations-node/register"
node app.js

Loading order matters in Node: if your application imports Express before the instrumentation is installed, Express will not be patched and you will see no spans. That is why the flag preloads it.

Try it
  1. Run the Flask service under opentelemetry-instrument and call both endpoints.
  2. In the Collector output find the span for /slow and locate its duration by comparing the start and end timestamps.
  3. Unset OTEL_SERVICE_NAME, restart, call again and find the new service name.
the duration is a little over 400 milliseconds, matching the sleep, and without the variable the service is reported as unknown_service (with the process name). A single missing environment variable is the most common reason a dashboard looks wrong.

Adding your own spans and attributes

Automatic instrumentation knows about HTTP requests and database calls. It knows nothing about your business: "validate the cart", "score this loan application", "run this model". To see inside your own code you create spans by hand. This is manual instrumentation, and it layers on top of the automatic kind.

The zero-code launcher has already set up the SDK, so your code needs only the API. A span is created with a tracer, obtained from the global provider, and used as a context manager so that it ends automatically:

app.py
from flask import Flask
from opentelemetry import trace
from opentelemetry.trace import Status, StatusCode

app = Flask(__name__)
tracer = trace.get_tracer("hello.orders", "1.0.0")


def price_for(item):
    prices = {"book": 12, "pen": 2}
    if item not in prices:
        raise KeyError(item)
    return prices[item]


@app.route("/order/<item>")
def order(item):
    with tracer.start_as_current_span("price-order") as span:
        span.set_attribute("order.item", item)
        span.add_event("lookup started")
        try:
            price = price_for(item)
        except KeyError as exc:
            span.record_exception(exc)
            span.set_status(Status(StatusCode.ERROR, "unknown item"))
            return {"error": "unknown item"}, 404
        span.set_attribute("order.price", price)
        return {"item": item, "price": price}

Run it again with opentelemetry-instrument python app.py and call curl localhost:8080/order/book and curl localhost:8080/order/ghost.

Several things are going on, and each one is a habit.

trace.get_tracer("hello.orders", "1.0.0") asks the global provider for a tracer and names the instrumentation scope. Use a stable, dotted name per module or component. start_as_current_span does three jobs at once: it creates the span, makes it the active span in the context, and, because it is a with block, ends it when the block exits, even if an exception flies out. Because the automatic Flask span is already active when order() runs, your new span becomes its child. You never set a parent. The context does it.

set_attribute adds a key-value pair. Prefer the standard semantic-convention names when one exists (http.request.method, db.system.name) and use your own dotted names (order.item) for business data. add_event records a timestamped note inside the span. record_exception attaches an exception, including its type, message and stack trace, as an event, and set_status marks the span as an ERROR so that backends can highlight and filter failed traces. Note that recording an exception does not by itself set the error status, so do both.

A warning about cardinality: an attribute with unbounded distinct values, such as a user ID or full URL with an ID in the path, is fine on a span but dangerous on a metric, where every distinct value creates a new time series. We return to that in the metrics section.

Name spans by operation, not by instance A good span name is low cardinality: price-order or GET /order/{item}, never price-order-book-7f3a. Put the variable parts in attributes. Backends group and aggregate by span name, and a name containing an ID turns one operation into a million.
Try it
  1. Add the manual span code, restart through opentelemetry-instrument, and call both order URLs.
  2. In the debug output find the price-order span for each call. Which one has an Events section with an exception?
  3. Check that the price-order span's Parent ID equals the ID of the Flask GET /order/... span.
the ghost request has an exception event and an error status, and your span is a child of Flask's. That parent link is the trace tree from earlier, built automatically.

Metrics and logs from your own code

Traces explain single requests. To count things over time, add a meter. As with spans, the API is what your code calls, and the SDK, started by the launcher, does the work.

metrics_demo.py
from opentelemetry import metrics

meter = metrics.get_meter("hello.orders", "1.0.0")

orders = meter.create_counter(
    "orders.processed", unit="{order}", description="Orders processed"
)
latency = meter.create_histogram(
    "orders.duration", unit="s", description="Time to price an order"
)

def record_order(item, seconds):
    orders.add(1, {"order.item": item})
    latency.record(seconds, {"order.item": item})

Each call to orders.add(1, {...}) adds one to the counter for that attribute combination. The unit uses the UCUM convention: s for seconds, {order} for a plain count of things. The histogram stores the distribution of durations, so a backend can compute percentiles.

The SDK does not send every measurement. A periodic metric reader collects the aggregated values and pushes them on an interval. The default is every 60 seconds, set by OTEL_METRIC_EXPORT_INTERVAL in milliseconds. During a demo that feels like nothing is happening, so shorten it:

BASH
export OTEL_METRIC_EXPORT_INTERVAL=5000

Two more ideas belong in your vocabulary. Temporality describes whether each exported number is cumulative (the total since the process started, the Prometheus style) or delta (the change since the last export). Cardinality is the number of distinct attribute combinations a metric has. A counter with attribute order.item and ten items has ten series. The same counter with user.id and a million users has a million, which can overwhelm a backend and your bill. Keep metric attributes to small, bounded sets such as item category, HTTP method or status class. The SDK's Views can rename, filter or drop attributes of a metric when a library exposes too many, which Mid-level covers.

Logs from your own code

For logs, the goal is correlation, not a new logging style. Keep using your language's standard logging. A bridge attaches OpenTelemetry to it so each record carries the active trace and span IDs. Logs are the least mature signal in Python and JavaScript at the time of writing: the Python logs module is still marked unstable in its import path (opentelemetry.sdk._logs), and the old LoggingHandler in the SDK is deprecated in favour of the opentelemetry-instrumentation-logging package. For a beginner the correct advice is modest: get traces and metrics working first, then add log correlation as a second step, and expect some API movement. In Java, Go and .NET the logs signal is stable or close to it.

Events are logs now Older material talks about a separate "Events API". It was removed from the specification. A standalone event is now simply a log record with an event name. The "events" you add to a span with add_event still exist for now, but the project has an accepted plan to deprecate the span event API too, so new code is steered toward emitting events through the Logs API.
Try it
  1. Call record_order from your /order/<item> handler and set OTEL_METRIC_EXPORT_INTERVAL=5000.
  2. Hit the endpoint a few times with two different items.
  3. Look in the Collector output for a Metric block with the name orders.processed and count its data points.
one data point per distinct order.item value, each with a running total. Count the points and you have counted the metric's cardinality.

Context propagation and sampling in practice

You met propagation in theory. In practice you rarely write code for it, because the instrumentation libraries do it: an instrumented HTTP client injects traceparent into outgoing requests, and an instrumented server extracts it from incoming ones. What you must do is check that each hop is instrumented and that nothing strips the headers.

You can prove propagation works with two small services. Write a second Flask app on port 8081 whose route calls http://localhost:8080/hello/aya using the requests library (install the matching opentelemetry-instrumentation-requests through opentelemetry-bootstrap, then run both under opentelemetry-instrument with different OTEL_SERVICE_NAME values). Call the second service and inspect the Collector output. The spans from both services carry the same trace ID, and the span of service A's outgoing call is the parent of service B's incoming span. Two processes, one trace.

BASH
curl -v http://localhost:8080/hello/aya 2>&1 | grep -i traceparent

Run that against a service that calls you and log incoming headers, and you will see the traceparent value the client attached. The default propagator list is tracecontext,baggage, set via OTEL_PROPAGATORS. Other formats such as b3 exist for systems that predate W3C Trace Context. The Jaeger and OT Trace propagators were deprecated in the specification, so prefer W3C for anything new.

Sampling: keeping only some traces

Recording every trace of a busy system is often too expensive. Sampling decides which traces to keep. There are two places to decide.

Head sampling happens in the SDK when a span starts, before anything is known about how the request will turn out. It is cheap and predictable. The default sampler is parentbased_always_on: record everything, but always follow the parent's decision, so a trace is either entirely kept or entirely dropped across services. To keep roughly ten percent of new traces:

BASH
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.1

The parentbased_ prefix is important. It means "if a parent already decided, obey it; otherwise apply the ratio at the root". Without it, every service would roll its own dice and your traces would come back with holes.

Tail sampling happens later, in the Collector, after the whole trace has been seen, so it can decide "keep every trace that contains an error or took longer than two seconds". That is powerful, and it is also stateful and memory-hungry, so it is a Senior topic. For now the rule is simple: head sampling is easy and blind, tail sampling is smart and costly, and on your laptop you need neither.

Sampling before you measure hides the problem A common beginner move is to set a one percent sampler on day one "to save money", then wonder why the rare slow request never appears in any trace. Start with everything on in development, measure volume, and add sampling deliberately when there is a bill or a load problem to solve.
Try it
  1. Run two services as described, with OTEL_SERVICE_NAME set to service-a and service-b.
  2. Call A so that it calls B, and compare the Trace IDs in the Collector output.
  3. Set OTEL_TRACES_SAMPLER=always_off on A only, restart it, and repeat.
with both on, the trace IDs match and B's parent is A's client span. With A off, nothing reaches the Collector from A, and because B is parent-based it inherits the unsampled decision and is silent too.

Configuring the SDK with environment variables

Every OpenTelemetry SDK is configured by the same set of environment variables, defined in the specification, which is why you can move between Python, Java and Node.js and still know what to type. You have met several. Here is the working set a beginner needs, grouped by purpose.

Purpose Variable Notes
Identity OTEL_SERVICE_NAME The most important one. Unset gives unknown_service
Identity OTEL_RESOURCE_ATTRIBUTES key=value,key2=value2
Destination OTEL_EXPORTER_OTLP_ENDPOINT Base URL. gRPC default http://localhost:4317, HTTP default http://localhost:4318
Destination OTEL_EXPORTER_OTLP_PROTOCOL grpc, http/protobuf or http/json
Destination OTEL_EXPORTER_OTLP_HEADERS key=value,key2=value2 for auth headers
Choice OTEL_TRACES_EXPORTER otlp, console, or none (same for metrics and logs; metrics also accepts prometheus)
Volume OTEL_TRACES_SAMPLER and _ARG Default parentbased_always_on
Propagation OTEL_PROPAGATORS Default tracecontext,baggage
Timing OTEL_METRIC_EXPORT_INTERVAL Milliseconds, default 60000
Switch OTEL_SDK_DISABLED true turns the SDK off entirely

Two details deserve attention. First, the difference between the base endpoint and the per-signal endpoint. With the HTTP protocol, OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4318 is a base: the SDK appends /v1/traces, /v1/metrics and /v1/logs itself. If instead you set OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, that value is used exactly as written, so it must include /v1/traces. Forgetting this gives a 404 from the Collector.

Second, the console exporter is the cheapest way to debug an SDK in isolation. Setting OTEL_TRACES_EXPORTER=console prints spans to the application's own standard output with no Collector involved, which tells you immediately whether the instrumentation is producing spans at all. If the console shows spans but the Collector shows nothing, the problem is the connection. If neither does, the problem is in the instrumentation.

BASH
export OTEL_TRACES_EXPORTER=console
opentelemetry-instrument python app.py

Java turns every variable into a system property by lowercasing it and swapping underscores for dots, so OTEL_SERVICE_NAME becomes -Dotel.service.name. You will also meet a newer approach called declarative configuration, a YAML file pointed at by OTEL_CONFIG_FILE. When it is set the SDK ignores the other OTEL_* variables except ones you reference explicitly inside the file. Support differs by language and it is beyond this level, but be aware that the old name OTEL_EXPERIMENTAL_CONFIG_FILE is deprecated.

Try it
  1. Run your service with OTEL_TRACES_EXPORTER=console and call an endpoint.
  2. Switch back to otlp with the protocol set to http/protobuf and the endpoint to http://localhost:4318.
  3. Set OTEL_EXPORTER_OTLP_TRACES_ENDPOINT=http://localhost:4318 (without the path) and observe the failure.
step one prints spans locally, step two works through the HTTP listener, and step three fails with a 404 because a per-signal endpoint is not given the /v1/traces path automatically.

Everyday Collector commands and configuration

You now know enough to work with the Collector day to day. This section collects the commands and config habits grouped by what you are trying to do.

Start, check and inspect

BASH
otelcol-contrib --config=config.yaml
otelcol-contrib validate --config=config.yaml
otelcol-contrib components
otelcol-contrib print-config --config=config.yaml
otelcol-contrib featuregate

print-config shows the final configuration after all merging and environment expansion, which is invaluable when config comes from several places. By default it hides secrets (--mode=redacted); --mode=unredacted shows them, so be careful where you paste the output. Older tutorials and even some docs pages tell you to add --feature-gates=otelcol.printInitialConfig. That gate was removed, and passing it now fails with no such feature gate "otelcol.printInitialConfig". Just run print-config. featuregate lists feature gates, the switches that control experimental and migrating behaviour.

Multiple configs and overrides

--config may be given more than once, and later files are merged over earlier ones. It accepts several URI schemes: file: (or a bare path), env:, yaml:, http: and https:. A classic pattern is one base file plus a small environment-specific override:

BASH
otelcol-contrib --config=base.yaml --config=prod-overrides.yaml
otelcol-contrib --config=config.yaml --set exporters::debug::verbosity=basic

--set overrides a single key after merging, using :: to separate nesting levels. A trap: if a later file contains an empty section such as a bare processors:, it is a null value and can wipe out what an earlier file defined. Write processors: {} or omit the key.

Environment variables in config

Inside the YAML you can read environment variables with ${env:NAME} and give a default with ${env:NAME:-fallback}:

config.yaml
exporters:
  otlp_http/vendor:
    endpoint: ${env:BACKEND_URL:-https://otlp.example.com}
    headers:
      authorization: Bearer ${env:BACKEND_TOKEN}

Never paste a real token into the file. Read it from the environment and let your orchestrator supply it. Note that the bare $NAME form is no longer expanded, and a literal dollar sign must be written $$.

Adding a real destination

The otlp_grpc exporter sends OTLP over gRPC to any OTLP-speaking backend, and otlp_http does the same over HTTP. These are the current names. Earlier versions called them otlp and otlphttp; those still work as deprecated aliases and the Collector logs a warning like "otlp" alias is deprecated; use "otlp_grpc" instead at startup. The trap to remember is that the receiver is still called otlp, while only the exporter was renamed.

config.yaml
exporters:
  otlp_grpc/jaeger:
    endpoint: jaeger:4317
    tls:
      insecure: true

Jaeger accepts OTLP natively on 4317 and 4318, so an otlp_grpc exporter pointed at it is all you need. The old dedicated Jaeger exporter is gone. tls.insecure: true disables TLS for a plaintext local connection, which is right for a demo on your laptop and wrong across the internet. Add the exporter to the traces pipeline's exporters list ([debug, otlp_grpc/jaeger]) to fan out to both.

For metrics the prometheus exporter exposes an endpoint that a Prometheus server can scrape:

config.yaml
exporters:
  prometheus:
    endpoint: 0.0.0.0:8889

Publish port 8889 from your container, add prometheus to the metrics pipeline, and scraping http://localhost:8889/metrics returns your counters and histograms in Prometheus format. The Prometheus guide covers what to do with them next.

Watching the Collector itself

The Collector reports on itself. Its own metrics are served on port 8888 in Prometheus format at http://localhost:8888/metrics, bound to localhost by default. The ones a beginner should learn to spot are otelcol_receiver_accepted_spans (data arrived), otelcol_exporter_sent_spans (data left), and the failure counters otelcol_receiver_refused_* and otelcol_exporter_send_failed_*. If accepted is rising and sent is not, the problem is on the export side.

Try it
  1. Run print-config on your file, then again with --set exporters::debug::verbosity=basic added, and compare.
  2. Add a prometheus exporter on 8889 to the metrics pipeline, publish the port, and curl /metrics after sending some metrics.
  3. Curl localhost:8888/metrics (publish that port too) and find otelcol_receiver_accepted_spans.
the override appears in the printed config, your orders_processed style metric is visible on 8889, and the accepted-spans counter equals the number of spans you generated. You have watched data enter and leave.

Sending traces to a real viewer

Reading spans in a terminal teaches the model, but you will want a picture. The lowest-friction way is to run a local Jaeger, an open-source trace viewer, and add it as a second exporter. Start it in Docker, on a shared network with your Collector so that the name jaeger resolves:

BASH
docker network create otel-net
docker run -d --rm --name jaeger --network otel-net \
  -p 16686:16686 jaegertracing/all-in-one

Then restart the Collector on the same network with the Jaeger exporter from the previous section in the traces pipeline:

BASH
docker run --rm --network otel-net \
  -p 127.0.0.1:4317:4317 -p 127.0.0.1:4318:4318 \
  -v "$(pwd)/config.yaml:/etc/otelcol-contrib/config.yaml" \
  otel/opentelemetry-collector-contrib:0.162.0

Open http://localhost:16686, choose the service hello-api, and click "Find Traces". You see the waterfall diagram from the start of this guide, now drawn for your own service, with the /slow request's 400 milliseconds as the long bar. Click a span to see its attributes. Nothing in your application changed to make this appear. You added an exporter to the Collector's config. That is the decoupling in action.

The same Collector could send the same spans to Grafana Tempo, to Elasticsearch, or to a commercial vendor by adding another exporter. The Langfuse guide and the Arize Phoenix guide are good next reads if your telemetry comes from LLM applications, since both can ingest OTLP.

Keep the debug exporter while you learn Leave debug in the exporters list next to the real backend. When the viewer shows nothing, one look at the Collector log tells you whether the data ever arrived. That splits every "nothing shows up" problem cleanly into before-the-Collector and after-the-Collector.
Try it
  1. Start Jaeger and the Collector on the same network, then call /slow and /order/ghost.
  2. In Jaeger find the trace for /order/ghost and look for the error marker on the price-order span.
  3. Open the exception event on that span and read the stack trace.
the failed span is flagged and carries the exception you recorded. This is the experience OpenTelemetry exists to give you, and you built it with a dozen lines of config.

Common errors and how to read them

Nearly every first-week problem falls into a small set. Learn to read these messages and you will fix most of them in a minute. Collector errors arrive with a prefix: failed to get config: means the file could not be parsed or resolved, and invalid configuration: means it parsed but does not make sense.

A reference to something undefined. service::pipelines::traces: references exporter "otlp_grpc/tempo" which is not configured. The pipeline names an ID that has no definition at the top of the file. Usually it is a typo in type/name, or you deleted the definition but not the reference. The same message exists for receivers, processors and extensions.

An unknown component type. unknown type: "file_log" for id: "file_log" (valid values: [...]). The component is not compiled into your binary, or its name is misspelled. The usual cause is using the core distribution for something that only exists in contrib. Run components and use the contrib build.

Invalid keys. decoding failed due to the following error(s): ... has invalid keys: flush. A key is misspelled, in the wrong place, or was removed in a newer release. The Collector is 0.x, so tutorials go stale. Check the component's README for the version you run, and check your indentation, which is also the usual cause.

Incomplete memory_limiter. 'check_interval' must be greater than zero or 'limit_mib' or 'limit_percentage' must be greater than zero. The limiter has no usable defaults. Set both, as in the config above.

Nothing pipeline-shaped. service must have at least one pipeline, must have at least one receiver, or must have at least one exporter. Your service block is empty or a pipeline is missing a list.

The removed logging exporter. the logging exporter has been deprecated, use the debug exporter instead. Replace the name.

Address already in use. listen tcp 0.0.0.0:4317: bind: address already in use. Another Collector, a Jaeger container, or an old process holds the port. Find and stop it, or change the endpoint.

A deprecation warning. "otlp" alias is deprecated; use "otlp_grpc" instead. This is a warning, not an error. Your config works. Rename it when convenient. You will see the same pattern for other components renamed to snake_case, such as filelog to file_log and k8sattributes to k8s_attributes.

Connection refused, silently. The Collector looks healthy, your app cannot connect, and nothing is logged. The receiver is bound to localhost inside a container. Set 0.0.0.0:4317.

Protocol and port mismatch. The application logs Transient error StatusCode.UNAVAILABLE encountered while exporting traces to localhost:4317, retrying in 1.00s, or the Collector sees HTTP/2 framing errors, or transport is closing. Sending gRPC to 4318 or HTTP to 4317 fails. Match the protocol to the port, and remember the Python distro defaults to gRPC.

TLS errors on export. tls: first record does not look like a TLS handshake means the exporter is trying TLS against a plaintext server: set tls: { insecure: true } for a local plaintext backend. x509: certificate signed by unknown authority is the reverse, and needs tls.ca_file pointing at the right certificate authority.

A 404 over HTTP. The per-signal endpoint lacks /v1/traces, or /v1/traces was appended to a gRPC endpoint. See the endpoint rule above.

The service name unknown_service. OTEL_SERVICE_NAME is not set in the process that actually runs your app. Check the environment of the process, not your shell.

No Python spans under a pre-fork server. With gunicorn --workers 2 the SDK's background threads do not survive the fork, so telemetry silently vanishes. Initialise the SDK in a post-fork hook. Similarly, Flask's debug=True reloader breaks instrumentation; use use_reloader=False. On slim Docker images, pip builds for instrumentation packages may fail for lack of a compiler; install gcc or build-essential and Python headers.

A Node service with no spans. The SDK loaded after the libraries. Preload it with --require or --import, and do not combine both mechanisms.

Read from the bottom, and read the whole line Collector errors are nested: the outermost wrapper is generic and the last clause names the real problem. Start from the end of the message, find the component ID and the key, and only then look at the top. When the message mentions a path like service::pipelines::traces, that is literally the location in your YAML.
Try it
  1. Delete the check_interval line from memory_limiter and run validate.
  2. Change the receiver endpoint to localhost:4317, run in Docker, and try telemetrygen from the host.
  3. Point an SDK set to http/protobuf at port 4317 and read the error in the app log.
you meet three of the most common messages on purpose. Seeing them once, with a known cause, makes them dull when they appear at two in the morning.

Putting it all together

Let us finish with one small end-to-end project that uses everything above. You will build an "orders" service, instrument it both automatically and by hand, run a Collector with a real viewer and a Prometheus endpoint, and confirm the whole path. Think of it as a checklist you can redo in fifteen minutes on any laptop.

  1. Make a project folder with app.py (the Flask service with the manual price-order span and the orders.processed counter), config.yaml (the Collector config), and a virtual environment with flask, opentelemetry-distro and opentelemetry-exporter-otlp installed, followed by opentelemetry-bootstrap -a install.

  2. Write the Collector config with the otlp receiver on both protocols at 0.0.0.0, the memory_limiter and batch processors, three exporters (debug, otlp_grpc/jaeger, prometheus), the health_check extension, and pipelines that send traces to [debug, otlp_grpc/jaeger] and metrics to [debug, prometheus].

config.yaml
receivers:
  otlp:
    protocols:
      grpc: { endpoint: 0.0.0.0:4317 }
      http: { endpoint: 0.0.0.0:4318 }

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
    spike_limit_percentage: 25
  batch: {}

exporters:
  debug: { verbosity: basic }
  otlp_grpc/jaeger:
    endpoint: jaeger:4317
    tls: { insecure: true }
  prometheus:
    endpoint: 0.0.0.0:8889

extensions:
  health_check: { endpoint: 0.0.0.0:13133 }

service:
  extensions: [health_check]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [debug, otlp_grpc/jaeger]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [debug, prometheus]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [debug]
  1. Validate, then run Jaeger and the Collector on a shared Docker network, publishing ports 4317, 4318, 8889 and 13133:
BASH
otelcol-contrib validate --config=config.yaml
docker network create otel-net
docker run -d --rm --name jaeger --network otel-net -p 16686:16686 jaegertracing/all-in-one
docker run -d --rm --name otelcol --network otel-net \
  -p 127.0.0.1:4317:4317 -p 127.0.0.1:4318:4318 \
  -p 127.0.0.1:8889:8889 -p 127.0.0.1:13133:13133 \
  -v "$(pwd)/config.yaml:/etc/otelcol-contrib/config.yaml" \
  otel/opentelemetry-collector-contrib:0.162.0
curl http://localhost:13133/
  1. Run the service with its identity and a short metric interval:
BASH
export OTEL_SERVICE_NAME=orders
export OTEL_RESOURCE_ATTRIBUTES=service.version=0.1.0,deployment.environment.name=dev
export OTEL_METRIC_EXPORT_INTERVAL=5000
export OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317
opentelemetry-instrument python app.py
  1. Generate traffic, including a failure:
BASH
for i in 1 2 3; do curl -s localhost:8080/order/book; done
curl -s localhost:8080/order/ghost
curl -s localhost:8080/slow
  1. Check each stage of the path. Read docker logs otelcol for the debug summaries (data arrived). Open http://localhost:16686 and find the failed orders trace with its exception event (traces flow). Curl http://localhost:8889/metrics and find orders_processed (metrics flow; the Prometheus exporter rewrites the dotted name into Prometheus style). If a stage is empty, you now know which link of the chain to inspect, because you can test each one separately: console exporter for the SDK, debug exporter for the Collector, viewer for the backend.

The exact rewritten metric name and any suffixes depend on the Prometheus exporter's naming rules for your version, so read the /metrics output rather than expecting a particular string.

What you can now do, and what comes next

You can now explain, without notes, that OpenTelemetry is a standard and a toolkit and not a backend; that the data flows from your instrumented code through an SDK and OTLP to a Collector and on to one or more backends; and that traces, metrics and logs are three views of one system joined by trace IDs and resource attributes. You can install and validate a Collector, read its config as receivers, processors and exporters wired into pipelines, instrument a Python service with no code changes, add your own spans with attributes, events and error status, add a counter and a histogram, set the SDK up with environment variables, and read the dozen error messages that account for most first-week pain.

Just as valuable are the habits you have picked up: set OTEL_SERVICE_NAME first, match protocol to port, keep the debug exporter close, put memory_limiter first and batch last, pin the Collector version, never paste secrets into config, and keep metric attributes low-cardinality.

What the Mid-level guide adds is the machinery behind these behaviours and the patterns that hold up in production: how processors, queues and retries actually behave, running the Collector as an agent and gateway in Kubernetes, enriching telemetry with Kubernetes metadata, filtering and transforming with the OpenTelemetry Transformation Language, controlling cardinality with Views, and testing your instrumentation. The Senior guide treats the Collector as a platform: tail sampling, scaling stateful components, multi-tenancy, security, upgrades, and when OpenTelemetry is not the right tool.

Natural neighbours in this catalogue are the Prometheus guide for the metrics side, the Kubernetes guide and the Helm guide for deploying Collectors, and the Langfuse guide and Arize Phoenix guide when your telemetry comes from LLM applications. One useful piece of news for that audience is that the gen_ai.* semantic conventions for LLM calls now live in their own repository, semantic-conventions-genai, separate from the main conventions, so look there rather than in older material.

Sources