Skip to content
Back to student guides
PrometheusDevOpsObservability3 levels127 sectionsCovers Prometheus 3.15

The Complete Prometheus Guide

Collect metrics and alert on them with Prometheus and PromQL. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
18sections
35examples

This is part one of three. It covers everything you need to do real work with Prometheus, not a teaser. By the end you can install a server, point it at a machine and at your own application, read the numbers it collects with its query language, turn a useful query into a saved rule, and write an alert that fires when something is genuinely wrong. Mid-level and Senior take the same topics further; nothing you learn here is thrown away.

The guide targets Prometheus 3.15, released in September 2026. Version 3.13 is the long-term-support line and is supported until July 2027, so if you work somewhere conservative you will meet it instead. Everything here works on both, and where a setting changed between 2.x and 3.x the text says so, because the internet is full of older tutorials that no longer match the tool. There is no supported 2.x release any more, so this guide teaches 3.x only.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and monitoring concepts only stick once you have watched a graph move because of something you did. You need a terminal, a laptop that can run a small program on port 9090, and curiosity. Nothing here needs a cloud account or a cluster.

What Prometheus is, and the problem it solves

Prometheus is an open-source monitoring and alerting toolkit. It regularly asks your programs "how are you doing right now?", writes the answers into a database built for numbers that change over time, lets you ask questions about that history, and tells a human when the answers look bad. It is a graduated project of the Cloud Native Computing Foundation, the same foundation that hosts Kubernetes, and it has become the default metrics system in the container world.

To see why it exists, picture running a web service without it. A customer writes in: "your site is slow." You log in to the server and run top. The CPU looks fine right now, but the slowness was twenty minutes ago and the evidence is gone. You add a log line, redeploy, and wait for it to happen again. Monitoring is the practice of recording numbers continuously, before you know what you will need to ask, so that when someone says "it was slow at 14:20" you can open a graph and look at 14:20.

Before Prometheus, teams typically had two styles of tool. Check-based systems such as Nagios ran a script on a schedule and showed a green or red light: disk is fine, disk is full. That answers "is it broken?" but not "how fast is it filling?" or "what did the last week look like?". Push-based metric systems such as StatsD and Graphite received numbers sent by your applications, but they made it hard to know whether silence meant "nothing happened" or "the application died". Prometheus combined the best ideas: store numeric time series like Graphite, but pull them so that a missing answer is itself a fact the system records.

TARGETSapps and exporters expose /metrics
→
PROMETHEUSscrapes, stores, evaluates rules
→
ALERTMANAGERgroups and routes alerts
→
A HUMANSlack, email, pager

The diagram is the whole architecture at beginner level. The word for the pulling step is scraping: every few seconds, Prometheus makes an ordinary HTTP request to an address such as http://myapp:8000/metrics and reads the plain text it gets back. That text is a list of metric names and their current values. Anything that can serve text over HTTP can be monitored, and that is why the ecosystem grew so quickly.

Three design decisions explain most of how Prometheus behaves, so notice them now.

Pull instead of push. Because Prometheus initiates every scrape, it always knows what it should be monitoring and can tell when a target stops answering. It records a synthetic metric called up with value 1 for a healthy scrape and 0 for a failed one. With a push system, a dead application simply stops sending and nobody notices. With Prometheus, death is a number you can alert on.

Every server stands alone. A Prometheus server keeps its data on its own local disk and does not depend on a cluster, a distributed database, or a message queue. When the network falls apart, the monitoring still works for everything it can reach. The price is that one server is one machine's worth of capacity and is not replicated by itself; the Senior guide covers how teams handle that.

Numbers, not events. Prometheus stores metrics: counts, rates, sizes, durations. It does not store individual log lines or individual request traces. It can tell you that 3 percent of requests failed in the last five minutes, but not which customer's request failed. For that you pair it with logs and traces, which other tools in this catalogue cover. It is also a poor fit where every single event must be counted exactly, such as billing, because a scrape sees a sampled picture at fixed intervals.

What people use it for:

🖥️

Infrastructure health

CPU, memory, disk and network for servers and virtual machines, through a small helper called node_exporter.

🌐

Service behaviour

Request rate, error rate and latency for your own APIs, the three numbers that describe most user-visible problems.

☸️

Kubernetes

The standard way to watch clusters: pods, nodes, deployments, all discovered automatically.

🤖

ML serving

Inference latency, throughput and GPU utilisation for model servers, which is why an MLOps engineer needs this tool.

Try it
  1. Think of a service you use or run. Write down three numbers that would tell you it is unhealthy.
  2. For each one, note whether it is a count of things that happened (requests, errors) or a level at a moment (memory in use, queue length).
  3. Note which of the three you could not reconstruct afterwards if you had not been recording it.
you will have a mix of counts and levels, and at least one number you would only have if you had been recording it all along. That split is the next concept, and the recording habit is the reason Prometheus exists.

The mental model: metrics, labels and time series

Prometheus has a small vocabulary, and once it is clear the rest of the tool is unsurprising. There are four nouns to hold: the metric, the label, the sample, and the time series they combine into.

A metric name says what is being measured: http_requests_total, node_memory_MemAvailable_bytes, process_cpu_seconds_total. By convention names use lowercase words joined by underscores, end in a unit when there is one (_seconds, _bytes), and counters end in _total. Since Prometheus 3.0 names may contain any UTF-8 characters, but the classic shape [a-zA-Z_:][a-zA-Z0-9_:]* is what nearly everything you meet will use. Colons are reserved for rules that you write yourself, which we meet later.

A label is a key and value pair that adds a dimension to a metric. The request counter by itself would tell you the total for the whole application. With labels it can tell you the total per method, per status code and per URL path. Labels are written inside curly braces:

TEXT
http_requests_total{method="GET", handler="/api/items", status="200"}

A sample is one measurement: a number together with the time it was taken, with millisecond precision. Values are 64-bit floating-point numbers, so even a count is stored as a float.

A time series is the stream of samples that share the same metric name and the same complete set of labels. This is the most important sentence in the guide. The metric name together with every label defines identity. http_requests_total{method="GET", status="200"} and http_requests_total{method="GET", status="500"} are two different time series, each with its own history, even though they share a name.

METRIC NAMEhttp_requests_total
+
LABELSmethod, handler, status
=
TIME SERIESone unique history
→
SAMPLES(timestamp, value) pairs

Internally the metric name is itself stored as a label called __name__. You will occasionally see it in queries such as {__name__="up"}, which is the same as writing up. Label names that begin with two underscores are reserved for Prometheus's own use, and an empty label value is treated as if the label did not exist at all.

Why labels need restraint

Labels are powerful, and that power has one sharp edge you should learn on day one. Every unique combination of label values is a separate time series, and each series costs memory and disk. The number of distinct series is called cardinality, and it is the main thing that makes a Prometheus server slow or expensive.

Suppose you label a request counter with method (about 5 values), status (about 10) and handler (about 50). That is at most 2,500 series, which is fine. Now someone adds a user_id label because it seems useful. With a million users the same metric becomes billions of potential series and the server runs out of memory. The rule is simple: label values must come from a small, bounded set. Never put user IDs, email addresses, request IDs, session tokens or raw URLs with query strings into labels. Those belong in logs.

The classic beginner trap A label with unbounded values is the most common way to hurt a Prometheus server. It works perfectly in testing with ten users and collapses in production with ten thousand. Before adding a label, ask: how many distinct values can this have a year from now? If the answer is "it grows with traffic", it is not a label.

Jobs and instances

Two labels appear on almost everything you collect, and Prometheus adds them for you. An instance is one single thing you scrape, identified by its host:port address. A job is a group of instances that do the same thing, such as all five copies of your API. So job="api", instance="10.0.0.12:8000" says "this number came from that copy of the API". Queries then ask either about one instance or about the whole job by summing across instances, and that is how you get from "one pod is slow" to "the service is slow".

For every scrape Prometheus also writes a few metrics of its own about the scrape itself. The one to remember is up: it is 1 when the last scrape succeeded and 0 when it failed. It is the simplest health check you will ever use, and the first query almost everyone runs.

Try it
  1. Write the full notation for a metric that counts disk reads on a server, with labels for the device and the host.
  2. Decide how many distinct series it produces if a host has 3 devices and you monitor 20 hosts.
  3. Now imagine you add a label holding the file name being read. Estimate how the count changes, and decide whether it belongs there.
60 series for the sensible version, and an unbounded count once file names appear. That contrast, small and bounded versus growing with data, is what cardinality means.

The four metric types

Your applications describe their numbers using four types. The type is a convention used by client libraries and documentation, and it tells you which query to use. Prometheus's database itself only stores floats (and, newer, native histograms), but picking the right query for the type is the difference between a correct graph and a nonsense one.

A counter is a number that only goes up, and resets to zero when the program restarts. Requests served, errors seen, bytes sent. A raw counter is almost useless to look at, because "41,872,390 requests since the last restart" tells you nothing. What you want is the rate of change: requests per second over the last five minutes. So the rule for counters is: never graph a counter directly; always wrap it in rate() or increase(). Both handle restarts correctly by noticing when the value drops and treating it as a reset.

A gauge is a number that can go up and down: memory in use, temperature, the number of items in a queue, the number of requests currently in flight. You can graph a gauge directly, and use functions such as avg_over_time or max_over_time to summarise it over a window.

A histogram measures how values are distributed, most often request durations. Instead of storing every duration, it counts how many observations fell into each of a set of buckets such as "under 0.1 seconds", "under 0.5 seconds", "under 1 second". A classic histogram called http_request_duration_seconds exposes three families of series: http_request_duration_seconds_bucket with an le label (less than or equal to) for each bucket boundary, http_request_duration_seconds_sum with the total of all observed values, and http_request_duration_seconds_count with the number of observations. The buckets are cumulative, and the function histogram_quantile() turns them into statements like "the 95th percentile latency is 240 milliseconds". Histograms can be combined across instances, which is why they are preferred.

A summary also measures distributions but calculates the percentiles inside the application and exposes them as quantile labels. It looks convenient and has a serious drawback: percentiles calculated separately in each instance cannot be averaged or combined into a correct overall percentile. For that reason new instrumentation nearly always uses histograms.

Type Goes Typical question Query habit
Counter only up how many per second? rate(x_total[5m])
Gauge up and down how much right now? use directly, or avg_over_time
Histogram buckets that only go up what is the 95th percentile? histogram_quantile(0.95, ...)
Summary precomputed quantiles what did this one instance see? read quantile labels, do not aggregate

Since version 3.9 there is also a stable feature called native histograms, which store a whole distribution in a single series with automatically sized buckets. They are cheaper and more precise than classic histograms, and you switch them on with the configuration option scrape_native_histograms: true. You do not need them to get started, and the Mid-level guide returns to them. The older command-line flag --enable-feature=native-histograms no longer does anything and only prints a warning, so ignore tutorials that still tell you to use it.

A naming shortcut You can usually guess the type from the name. A name ending in _total is a counter. A name ending in _bucket, _sum or _count belongs to a histogram. A name ending in a unit such as _bytes or _seconds with no other suffix is usually a gauge.
Try it
  1. Classify each of these as counter, gauge, histogram or summary: bytes received by a network card, free disk space, the time each checkout request takes, the number of logged-in users.
  2. For each counter, write the query shape you would use to turn it into something readable.
counter, gauge, histogram, gauge. Only the first needs rate(), the third needs histogram_quantile(), and the two gauges can be graphed as they are.

Installing Prometheus and checking the setup

Prometheus is a single statically linked program written in Go, so installing it means getting one file and running it. That is also why the binary download is the best way to learn: nothing is hidden. Pick the path that matches your machine, then run the same verification steps.

Linux

Download the release archive from the official GitHub releases page, check its checksum, and unpack it. Use linux-arm64 instead of linux-amd64 on an ARM machine.

BASH
VER=3.15.0
curl -LO https://github.com/prometheus/prometheus/releases/download/v${VER}/prometheus-${VER}.linux-amd64.tar.gz
curl -LO https://github.com/prometheus/prometheus/releases/download/v${VER}/sha256sums.txt
sha256sum --check --ignore-missing sha256sums.txt
tar xzf prometheus-${VER}.linux-amd64.tar.gz
cd prometheus-${VER}.linux-amd64
ls

Inside you will find the prometheus server, a helper called promtool that we use for checking files, an example prometheus.yml, and the licence files. The checksum step matters: it confirms the file you downloaded is the file the project published. sha256sum should print prometheus-3.15.0.linux-amd64.tar.gz: OK.

macOS

The simplest route is Homebrew, which installs both the server and promtool:

BASH
brew install prometheus
prometheus --version

Homebrew places the example config at $(brew --prefix)/etc/prometheus.yml. You can run it in the foreground while learning with prometheus --config.file=$(brew --prefix)/etc/prometheus.yml, or as a background service with brew services start prometheus. If you prefer the tarball, download prometheus-3.15.0.darwin-arm64.tar.gz on Apple Silicon or the darwin-amd64 build on Intel, and if macOS refuses to open the unsigned binaries, remove the quarantine flag with xattr -d com.apple.quarantine ./prometheus ./promtool.

Windows

Download prometheus-3.15.0.windows-amd64.zip from the releases page, extract it, and run it from PowerShell inside the extracted folder with .\prometheus.exe --config.file=prometheus.yml. Allow port 9090 through Windows Defender Firewall if it asks. The project does not ship a Windows service wrapper, so for anything permanent people use Docker, WSL2, or a third-party wrapper such as NSSM.

Docker, on any system

If you already have Docker (see the Docker guide), one command starts a server:

BASH
docker run -d --name prometheus -p 9090:9090 \
  -v "$PWD/prometheus.yml":/etc/prometheus/prometheus.yml \
  -v prometheus-data:/prometheus \
  prom/prometheus:v3.15.0

The first volume mounts your config file into the place the image expects it. The second is a named volume for the database, and you should never skip it: without it the collected data vanishes with the container. One detail catches people out. The image starts with default arguments that point at /etc/prometheus/prometheus.yml and /prometheus. If you pass any extra arguments, they replace those defaults, so you must repeat --config.file and --storage.tsdb.path yourself.

Kubernetes, in one paragraph

On Kubernetes nobody installs the binary by hand. The usual route is the Prometheus Operator and the kube-prometheus-stack Helm chart from the prometheus-community repository, which installs a server, Alertmanager, node metrics and ready-made dashboards together. That belongs to the later guides. Learn on a laptop first, because the concepts are identical and the moving parts are fewer. The Helm guide is the natural follow-up when you get there.

Verify the install

Whatever route you took, verify it the same way. Start the server with the example configuration that came with it:

BASH
./prometheus --config.file=prometheus.yml

The server runs in the foreground and prints log lines. The two you want are msg="Completed loading of configuration file" and msg="Server is ready to receive web requests.". Open a second terminal and ask it three things:

BASH
./prometheus --version
curl -s localhost:9090/-/healthy
curl -s localhost:9090/-/ready

--version prints the version, which should say 3.15.0. The /-/healthy endpoint answers "Prometheus Server is Healthy." as soon as the process is alive, and /-/ready answers "Prometheus Server is Ready." once it has finished starting and can serve queries. During a slow start the ready endpoint returns 503 "Service Unavailable", which is normal and is how orchestrators know when to send traffic. Finally, open http://localhost:9090 in a browser. You should see the query page of the Prometheus interface.

Ports to recognise Prometheus listens on 9090. Its companions have their own well-known defaults: Alertmanager on 9093, the Pushgateway on 9091, node_exporter on 9100, and blackbox_exporter on 9115. You will see these numbers constantly, and recognising them makes configuration files readable.
Try it
  1. Install Prometheus using whichever route fits your machine and start it with the example config.
  2. Run the three curl and version commands above and confirm each answers as described.
  3. Open http://localhost:9090, type up into the query box and press Execute.
one row with job="prometheus", instance="localhost:9090" and the value 1. Prometheus is already monitoring itself, and that single row is a complete working monitoring system.

The configuration file, line by line

Prometheus is configured with one YAML file, usually called prometheus.yml, whose path you give with --config.file. Nearly every question of the form "how do I make Prometheus monitor X" is answered by editing this file, so it is worth reading slowly once. Here is the example that ships with the project, which is also the smallest useful configuration:

prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

alerting:
  alertmanagers:
    - static_configs:
        - targets:
          # - alertmanager:9093

rule_files:
  # - "first_rules.yml"

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

The file has four blocks that matter at this level.

global holds defaults for everything else. scrape_interval is how often Prometheus scrapes each target, and evaluation_interval is how often it evaluates your rules. One detail trips people up: the built-in default for both is one minute, not fifteen seconds. The example sets 15 seconds, which is a common and sensible choice because it gives a graph a fresh point four times a minute without loading the targets. There is also scrape_timeout, the time Prometheus waits for one answer, which defaults to 10 seconds and must never be larger than the interval.

scrape_configs is the heart of the file: a list of jobs. Each entry has a job_name (which becomes the job label) and a list of targets, here hard-coded under static_configs. A target is just host:port. By default Prometheus fetches the path /metrics over HTTP, and you can change that with metrics_path and scheme if your application differs.

rule_files lists files of recording and alerting rules, which come later in this guide. alerting says where Alertmanager lives. Both are empty in the example, and the commented lines show the shape to fill in.

There are many more keys in the full reference, for authentication, TLS, relabeling, limits and dozens of ways to discover targets automatically. You will meet a few in later sections, and the rest when a real need appears. The important skill now is not memorising keys but checking a file before you use it.

Validate before you run

Prometheus ships with promtool, a checker that understands the configuration syntax. Make it a reflex to run it after every edit:

BASH
./promtool check config prometheus.yml

On success it prints SUCCESS: prometheus.yml is valid prometheus config file syntax. On a mistake it tells you the line and the problem, for example an unknown field name or a bad indentation. This matters more than it might seem, because a server that is told to reload a broken file keeps running the old configuration and logs an error, so a typo can silently leave you with stale settings.

YAML is strict about spaces YAML uses indentation to express structure, and it does not allow tab characters. Two spaces per level is the convention. If promtool reports a parsing error near a line you did not touch, look at the indentation of the line above it. Copying examples from a web page is the usual source of mixed tabs and spaces.

Changing the configuration while running

After editing the file you do not need to stop the server. There are three ways to tell it to re-read the file. You can send the process a hang-up signal with kill -HUP <pid>, which always works. You can enable the lifecycle endpoints by starting with --web.enable-lifecycle and then call curl -X POST http://localhost:9090/-/reload. Or you can start with --config.auto-reload, a stable option since version 3.12, and Prometheus will check the file every 30 seconds by itself. Whichever you choose, confirm the result afterwards by querying prometheus_config_last_reload_successful: it is 1 when the last reload worked and 0 when the new file was rejected.

The lifecycle flag deserves a caution. It is off by default because it lets anyone who can reach the port reload or shut down your server. On a laptop that is fine, and on a shared network it is not.

Try it
  1. Copy the example config, then introduce a deliberate mistake by indenting job_name with a tab or removing a colon.
  2. Run ./promtool check config on it and read the message.
  3. Fix the file, then change scrape_interval to 5s, start the server with --web.enable-lifecycle, and reload it with the curl command.
first an error naming the broken line, then SUCCESS after the fix, and after the reload a log line saying the configuration file was loaded. Breaking things safely on purpose is the fastest way to learn what the error messages mean.

A tour of the web interface

The interface at port 9090 is where you will spend the first hours, and it is deliberately simple. Version 3.0 replaced the old interface with a new one, so screenshots in older tutorials will look different, but the same ideas are there.

The Query page is the main one. You type an expression into the box and press Execute. The result appears as a Table, showing the current value of each matching series, or as a Graph, showing how each value changed over a time window you choose. Start with the table, because it shows you exactly which series match and which labels they carry. Then switch to the graph when you want to see change over time. As you type, the box suggests metric names, which is the best way to discover what a server actually holds.

The Status menu has the pages you use for diagnosis. Target health lists every scrape target, with its state (UP or DOWN), its labels, how long ago it was last scraped, how long the scrape took, and the last error message if there was one. This is where you look whenever a graph is empty. Service discovery shows what each discovery mechanism found, and what labels the targets had before and after relabeling. Runtime and build information shows the version and start time. Configuration shows the configuration that is actually loaded, which is useful when you are not sure that your edit was picked up. TSDB status lists the metrics and labels with the most series, which is how you find a cardinality problem. And Flags shows the command-line settings the server started with.

The Alerts page lists every alerting rule and its state. You will use it once you write rules. Together, these pages make Prometheus unusually transparent: when something looks wrong, there is almost always a page that shows the state directly, before you have to read logs.

Empty graph? Check three places in order First, Status then Target health: is the target UP? Second, the Table view: does the metric name exist at all, and do your label matchers match its labels exactly? Third, the time range: a metric that only started existing five minutes ago will not show in a six-hour window's early part. Nine empty-graph problems out of ten are one of these.
Try it
  1. Open Status, then Target health, and read every column for the prometheus job.
  2. On the Query page, type prometheus_ and scroll through the suggestions.
  3. Query prometheus_tsdb_head_series and view it as a graph.
a target that is UP with a scrape duration of a few milliseconds, several hundred metric names that start with prometheus_, and a line showing how many series the server is holding in memory right now.

Monitoring a real machine with node_exporter

Watching Prometheus watch itself proves the plumbing, but it is not monitoring anything you care about. Most software does not speak the Prometheus format natively. Linux does not expose a /metrics page; it exposes files under /proc and /sys. The bridge is an exporter: a small separate program that reads the statistics of some other system and serves them in the Prometheus text format. The exporter for Linux and other Unix-like machines is node_exporter, and it is the first one nearly everyone installs. (On Windows the equivalent is windows_exporter, which listens on port 9182.)

The pattern generalises: there are exporters for databases such as MySQL, PostgreSQL and Redis, for message queues, for network devices, for GPUs, and for probing websites from the outside (the blackbox exporter). The shape is always the same. Run the exporter next to the thing it reports on, then add it to scrape_configs as one more job.

Install node_exporter from the Prometheus download page, or with your package manager. The version at the time of writing is 1.12. Running it needs no configuration:

BASH
tar xzf node_exporter-1.12.1.linux-amd64.tar.gz
cd node_exporter-1.12.1.linux-amd64
./node_exporter

It starts listening on port 9100. Before wiring it into Prometheus, look at what it serves, because reading the exposition format once removes all mystery from it:

BASH
curl -s localhost:9100/metrics | head -n 20

The output looks like this:

TEXT
# HELP node_cpu_seconds_total Seconds the CPUs spent in each mode.
# TYPE node_cpu_seconds_total counter
node_cpu_seconds_total{cpu="0",mode="idle"} 184523.48
node_cpu_seconds_total{cpu="0",mode="user"} 3921.17
# HELP node_memory_MemAvailable_bytes Memory information field MemAvailable_bytes.
# TYPE node_memory_MemAvailable_bytes gauge
node_memory_MemAvailable_bytes 5.123584e+09

Lines starting with # HELP describe the metric in words, lines starting with # TYPE give its type, and the other lines are the samples: a name, optional labels in braces, and a value. That is the entire format. Anything that can print text like this can be monitored, which is why a beginner can write a working exporter in twenty lines of any language.

Now add a second job to the configuration file, next to the existing one:

prometheus.yml
scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

  - job_name: "node"
    static_configs:
      - targets: ["localhost:9100"]
        labels:
          environment: "laptop"

The labels block under a target adds extra labels to every series from those targets, which is a simple way to mark an environment or a team. Validate the file, reload the server, and open Status then Target health. Both jobs should now show UP. Notice that you never listed a port for the metrics path or the scheme: the defaults of /metrics and http applied.

Exporters run on the machine, not in Prometheus A common confusion is to look for node_exporter inside the Prometheus installation. It is a separate program and a separate process. Prometheus only ever speaks HTTP to it. If the exporter is stopped, the job shows DOWN and up{job="node"} becomes 0. That is exactly the behaviour you want an alert to catch.
Try it
  1. Start node_exporter and read the first twenty lines of its /metrics output.
  2. Add the node job to prometheus.yml, validate it with promtool, and reload.
  3. Stop node_exporter with Ctrl-C, wait twenty seconds, and look at Target health again. Start it and look once more.
the target turns DOWN with a "connection refused" error within one or two scrape intervals, and returns to UP after you restart it. You have just watched the up metric do its job.

PromQL part one: selecting data

The data is now flowing. The skill that turns it into answers is PromQL, the Prometheus Query Language. It is a language designed for one job: selecting time series and calculating on them. It is not SQL and it does not look like it, so give yourself an hour to get used to the shape. The good news is that a small set of ideas covers most real use.

Type queries into the box on the Query page. The simplest query is just a metric name, which returns the current value of every series with that name:

PROMQL
node_memory_MemAvailable_bytes

This is called an instant vector: one value per matching series, all at the same moment. Most queries start here. To narrow the result you add label matchers in braces. There are four operators, and they are worth memorising:

Matcher Meaning Example
= label equals the value {job="node"}
!= label does not equal the value {mode!="idle"}
=~ label matches a regular expression {status=~"5.."}
!~ label does not match the expression {mountpoint!~"/run.*"}

Regular expressions in Prometheus are fully anchored, which means the pattern must match the whole label value, as if it had a caret at the start and a dollar sign at the end. So status=~"5.." matches 500 and 503 but not 2500. You do not write the anchors yourself. Combining matchers narrows the result further:

PROMQL
node_cpu_seconds_total{job="node", mode!="idle"}

One rule surprises beginners: a selector must contain at least one matcher that cannot match an empty string. The expression {job=~".*"} is refused because it could match everything, while {job=~".+"} is allowed. The easy way to stay clear of the error is to start every selector with a metric name.

Range vectors and the time window

An instant vector is one point per series. To look at a stretch of history you add a time range in square brackets, which turns it into a range vector:

PROMQL
node_cpu_seconds_total{mode="idle"}[5m]

This returns, for each series, all the samples from the last five minutes. Durations use the units ms, s, m, h, d, w and y. You rarely want to look at a range vector on its own, because it is a list of raw samples. Its job is to be handed to a function such as rate() that reduces it back to one number per series. That pairing, range vector in and instant vector out, is the pattern behind most useful queries.

You can also ask about the past directly with the offset modifier. node_memory_MemAvailable_bytes offset 1h returns the value one hour ago, which is handy for comparing "now" with "this time yesterday" using offset 1d. Since version 3.14 durations can even be small expressions, such as [5m * 2], although you will rarely need that at the start.

A 3.0 change that affects old examples Since Prometheus 3.0, range windows are left-open and right-closed: a sample exactly at the start of the window is excluded. In practice this only bites in tiny windows. If you read an old example that uses a range as short as the scrape interval and gets no data, widen the range.

What the four types of value are

Every PromQL expression returns one of four kinds of value, and error messages talk about them, so learn the names now. An instant vector is a set of series with one sample each. A range vector is a set of series with a window of samples each. A scalar is a single plain number such as 0.95. A string is rarely used. The most common beginner error message, expected type range vector in call to function "rate", got instant vector, simply means you forgot the square brackets.

Try it
  1. Query node_filesystem_avail_bytes and read the labels on each row in the Table view.
  2. Narrow it to one mount point with {mountpoint="/"}, then exclude temporary file systems with fstype!~"tmpfs|overlay".
  3. Add [10m] to the query and read what comes back.
a handful of rows that shrinks as you add matchers, and then, with the range added, a list of timestamped values per series. If you get a parse error, read it: it names the position of the problem.

PromQL part two: rates, aggregation and arithmetic

Selecting raw series only gets you so far. The real power arrives with three operations: rates for counters, aggregation across series, and arithmetic between them.

Rates for counters

Recall that counters only go up. The question you actually ask is "how fast is it going up?". The function rate() answers it: given a range vector, it returns the per-second average increase over that window, and it corrects for counter resets. To see how much CPU time your machine's CPUs spent per second in each mode:

PROMQL
rate(node_cpu_seconds_total[5m])

The five minutes in brackets is the window that rate() averages over. A good rule is to pick a window at least four times your scrape interval, so with 15-second scrapes five minutes is comfortable and one minute is the practical minimum. A window that is too short gives jumpy graphs and can return nothing when a scrape is missed. A window that is too long smooths away real spikes.

Two relatives are worth knowing. increase(x[1h]) gives the total increase over the window, which reads more naturally for questions such as "how many errors in the last hour". It is rate() multiplied by the window length. irate() uses only the last two samples and so is very twitchy; it is useful for zoomed-in graphs of fast-changing values and a poor choice for alerts. When in doubt, use rate().

Aggregation across series

The CPU query above returned a row for every CPU core and every mode, which is too much. Aggregation operators collapse many series into fewer. The main ones are sum, avg, min, max, count, and topk (the top N by value). You choose which labels survive with by, or which to discard with without:

PROMQL
sum by (mode) (rate(node_cpu_seconds_total[5m]))

This keeps only the mode label and sums across cores, giving one number per mode. The same idea scales from one laptop to a thousand pods: to get the request rate of a whole service you sum across its instances, and to see the busiest five you use topk(5, ...). Notice how the order reads from the inside out. The innermost part selects and applies rate(), and the outer part aggregates the result.

Rate first, then sum Always apply rate() to the individual counters, and sum the results afterwards. Writing rate(sum(x_total)[5m]), or summing counters before taking a rate, produces wrong numbers when one instance restarts, because the reset can no longer be seen inside the combined total. The correct shape is sum(rate(x_total[5m])).

Arithmetic and comparison

You can use ordinary operators, + - * / % ^, between a series and a number, or between two sets of series. Between two sets, Prometheus matches series that have identical labels and calculates pair by pair. That is how you turn two raw gauges into a percentage:

PROMQL
100 * (1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes)

This reads: the fraction of memory that is available, subtracted from one, expressed as a percentage, which is the percentage of memory in use. The two metrics carry the same labels, so they pair up automatically. A classic CPU utilisation query builds on the idle mode:

PROMQL
100 * (1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m])))

Read it inside out: the per-second idle time of each core over five minutes, averaged over the cores of each machine, turned into "how much of the time was not idle", as a percentage. Comparison operators such as >, < and == work as filters. node_filesystem_avail_bytes < 10 * 1024 * 1024 * 1024 keeps only the series below 10 GiB and drops the rest. That filtering behaviour is exactly what an alert needs: it fires for every series that survives the comparison, and stays silent when the result is empty.

When the labels on two sides do not match, you can tell Prometheus which labels to compare with on(...) or which to ignore with ignoring(...). The Mid-level guide covers matching across different label sets in depth. At this stage it is enough to know the error that signals it: many-to-many matching not allowed means that the two sides have more than one series per match group, and you need to aggregate or choose labels more carefully.

A short list of queries worth memorising

Question Query
Which targets are down? up == 0
How many targets per job? count by (job) (up)
Is a series missing entirely? absent(up{job="api"})
Requests per second per job sum by (job) (rate(http_requests_total[5m]))
Error ratio sum(rate(http_requests_total{status=~"5.."}[5m])) / sum(rate(http_requests_total[5m]))
95th percentile latency histogram_quantile(0.95, sum by (le) (rate(http_request_duration_seconds_bucket[5m])))
Disk full in four hours? predict_linear(node_filesystem_avail_bytes[6h], 4*3600) < 0

Two of these deserve a word. absent() returns 1 only when the series does not exist, so it catches the case that up == 0 cannot: a target that was never configured or has disappeared from discovery, leaving no up series to be zero. histogram_quantile() takes the bucket series of a classic histogram, and you must keep the le label when aggregating, because the buckets are the information it works from. The predict_linear() function fits a straight line through the window and extrapolates, which turns "the disk is 80 percent full" into the far more useful "the disk will be full before morning".

Try it
  1. Run rate(node_cpu_seconds_total[5m]) and count the rows, then run the sum by (mode) version and count again.
  2. Run the memory percentage query and compare the number with the output of free -m or your system monitor.
  3. Write a query that returns only the filesystems that are more than 80 percent full.
fewer rows after aggregation, a memory figure within a point or two of the operating system's, and a filter query that is often empty, which is the correct answer when nothing is wrong.

Recording rules: saving a query as a metric

Some queries are expensive, or are used in many places. A dashboard with twenty panels that each compute the same error ratio across thousands of series will make the server do the same work twenty times every refresh. A recording rule solves this by evaluating the expression on a schedule and saving the result as a new time series, which later queries read cheaply.

Rules live in separate files listed under rule_files. A rule file groups rules, and the rules in a group are evaluated in order, at the group's interval:

rules.yml
groups:
  - name: node_rules
    rules:
      - record: instance:node_cpu_utilisation:ratio_rate5m
        expr: 1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))
      - record: instance:node_memory_utilisation:ratio
        expr: 1 - node_memory_MemAvailable_bytes / node_memory_MemTotal_bytes

Each rule has a record (the name of the new metric) and an expr (the query). The naming style level:metric:operations is a convention worth following, and colons in a name are reserved for exactly this purpose. The first part says which labels are kept (here instance), the middle says what is measured, and the last says which operations were applied. It looks fussy, but a name like that tells a colleague what a series is without opening the rule file.

Hook the file into the configuration, then check both files before reloading:

prometheus.yml
rule_files:
  - "rules.yml"
BASH
./promtool check rules rules.yml
./promtool check config prometheus.yml

After a reload, wait one evaluation interval and query instance:node_cpu_utilisation:ratio_rate5m. It now behaves like any other metric. The Status menu also shows a Rule health page listing each group, how long its last evaluation took and whether it failed. Keep an eye on that duration: if a group takes longer than its interval, evaluations are skipped.

When a recording rule is worth it Record a query when it is used in several dashboards or alerts, when it aggregates a large number of series, or when you want a stable, named building block that other rules can use. Do not record every query you write; each recording rule creates new series that cost storage. A rule you can explain in one sentence and that something actually uses is a good rule.
Try it
  1. Create rules.yml with the two rules above, reference it from prometheus.yml, and check both with promtool.
  2. Reload the server and wait 30 seconds.
  3. Query the recorded metric names in the Table view.
one row per monitored machine with a value between 0 and 1. If the query returns nothing, open Status then Rule health and read the error for the group.

Alerting: from a query to a notification

A dashboard only helps when someone is looking at it. Alerting rules watch a query for you and raise an alert when it is true. Prometheus itself only decides that an alert is firing. Deciding who to tell and how is the job of a second program, Alertmanager, which keeps the two concerns cleanly apart.

Writing an alerting rule

Alerting rules sit in the same kind of file as recording rules. Add a new group to rules.yml:

rules.yml
groups:
  - name: node_alerts
    rules:
      - alert: InstanceDown
        expr: up == 0
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "{{ $labels.job }} target {{ $labels.instance }} is down"
          description: "The scrape has failed for more than 2 minutes."

      - alert: DiskWillFillSoon
        expr: predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4 * 3600) < 0
        for: 30m
        labels:
          severity: ticket
        annotations:
          summary: "Disk {{ $labels.mountpoint }} on {{ $labels.instance }} will fill within 4 hours"

An alerting rule has four parts. The alert field is its name. The expr is a query, and the alert fires for every series the expression returns, which is why comparisons such as up == 0 (a filter) fit so well. labels are extra labels attached to the alert, and the severity label is a convention that Alertmanager can use for routing. annotations carry human-readable text, and they can include templates such as {{ $labels.instance }}, which inserts the label value of the series that triggered it, and {{ $value }}, which inserts the number.

The pending state and for

The for clause is what keeps you from being woken by a blip. An alert moves through three states. It is inactive when the expression returns nothing. When the expression first returns something it becomes pending, and Prometheus starts a timer. Only when the expression has stayed true for the whole for duration does the alert become firing and get sent to Alertmanager. If the condition clears during the pending phase, the alert quietly goes back to inactive and nobody hears about it.

INACTIVEexpression is empty
→
PENDINGtrue, waiting out "for"
→
FIRINGsent to Alertmanager

Choosing for is a judgement call between speed and noise. Two minutes is reasonable for "a target is down", since a single failed scrape is usually a hiccup. Thirty minutes suits slow-moving problems like disk fill. Without a for value the alert fires on the first evaluation that is true, which is almost always too jumpy for paging a human.

Alertmanager: deduplicate, group, route

Alertmanager is a separate program, version 0.34 at the time of writing, that listens on port 9093. Prometheus sends it the firing alerts, and Alertmanager does the things you actually want from a notification system. It deduplicates repeated alerts, groups related ones so that fifty failing pods produce one message instead of fifty, routes alerts to different receivers based on their labels, lets you silence alerts during maintenance, and inhibits lower-priority alerts when a higher-priority one already explains them.

Its configuration is a routing tree. This minimal one groups by alert name and sends everything to a webhook:

alertmanager.yml
route:
  receiver: team-webhook
  group_by: [alertname]
  group_wait: 30s
  group_interval: 5m
  repeat_interval: 4h

receivers:
  - name: team-webhook
    webhook_configs:
      - url: "http://localhost:5001/"

group_wait is how long Alertmanager waits after the first alert of a group before sending, so that related alerts arrive together. group_interval is how long before it sends updates about a changed group, and repeat_interval is how often it reminds you about a problem that is still open. Real setups replace the webhook with Slack, email, PagerDuty or Microsoft Teams receivers, and use child routes to send severity: page alerts somewhere louder than severity: ticket ones.

Start it, then tell Prometheus where it is by filling in the alerting block. Since Prometheus 3.0 only version 2 of the Alertmanager API is supported, and it is the default:

BASH
./alertmanager --config.file=alertmanager.yml
prometheus.yml
alerting:
  alertmanagers:
    - static_configs:
        - targets: ["localhost:9093"]

Test a rule before trusting it

You do not have to break a production system to see whether an alert works. On your laptop, stop node_exporter and watch the Alerts page: InstanceDown becomes pending, and after two minutes firing. Open http://localhost:9093 and the alert appears in Alertmanager. For rules you need to be sure about, promtool test rules runs a rule against synthetic input series, so you can assert in a file that "with these values, this alert fires at minute five". That belongs in continuous integration, and the Mid-level guide shows it.

An alert nobody can act on is noise Every alert should answer: what is wrong, who should care, and what do they do first? Beginners tend to alert on causes (CPU above 80 percent) when users feel symptoms (slow responses, errors). A page at 3 a.m. for a busy CPU that hurts nobody teaches the team to ignore pages. Start with a few symptom-based alerts, such as target down, error ratio too high, and disk about to fill, and add more only when an incident shows the need.
Try it
  1. Add the InstanceDown rule, validate with promtool check rules, and reload.
  2. Start Alertmanager with the minimal config and point Prometheus at it.
  3. Stop node_exporter, watch the Alerts page for the pending state and then the firing state, then check localhost:9093.
the alert moves from inactive to pending to firing after the two-minute wait, then shows up in Alertmanager, and clears again once you restart the exporter.

Instrumenting your own application

Exporters cover other people's software. For your own services, you add instrumentation: a few lines of code using a client library that keeps counters and gauges in memory and serves them on /metrics. Official libraries exist for Go, Java, Python and Ruby, and the community maintains many more. The Python one, prometheus-client, is the quickest to try and the one most used in ML serving code. Install it with pip install prometheus-client.

The example below is a tiny program that pretends to serve requests. It exposes a counter of requests by outcome, a histogram of how long each one takes, and a gauge of requests currently being handled:

app.py
import random
import time

from prometheus_client import Counter, Gauge, Histogram, start_http_server

REQUESTS = Counter(
    "demo_requests_total", "Requests handled, by outcome.", ["outcome"]
)
LATENCY = Histogram(
    "demo_request_duration_seconds", "Time spent handling a request."
)
IN_FLIGHT = Gauge("demo_requests_in_flight", "Requests being handled now.")


def handle_request() -> None:
    with IN_FLIGHT.track_inprogress(), LATENCY.time():
        time.sleep(random.uniform(0.01, 0.4))
        outcome = "error" if random.random() < 0.05 else "ok"
        REQUESTS.labels(outcome=outcome).inc()


if __name__ == "__main__":
    start_http_server(8000)   # serves /metrics on port 8000
    while True:
        handle_request()

Run it with python app.py, then curl -s localhost:8000/metrics | grep demo_. You will see demo_requests_total{outcome="ok"} and demo_requests_total{outcome="error"} climbing, and the demo_request_duration_seconds_bucket family with its le labels. Notice how little code this took: declaring each metric once at module level, and updating it where the work happens. A few habits make instrumentation age well.

Declare metrics once, at the top. Create them as module-level objects, not inside the function that handles requests. Creating the same metric twice raises an error, and it would also reset the values.

Follow the naming conventions. Use a prefix for your application, end counters in _total, and put the unit in the name in base units: seconds rather than milliseconds, bytes rather than megabytes. Dashboards and alerts written by other people assume this.

Keep label values bounded. The outcome label has two values, which is ideal. Never label with anything that grows with users or requests, for the reasons in the cardinality section.

Measure the things users feel. The three most useful signals for a service are the rate of requests, the errors among them, and the duration of each. Many teams call this the RED method, and this single example already provides all three.

Then add the application as a scrape job and watch its numbers appear:

prometheus.yml
scrape_configs:
  - job_name: "demo-app"
    static_configs:
      - targets: ["localhost:8000"]

Now the queries from earlier apply to your own code. The request rate is sum(rate(demo_requests_total[5m])), the error ratio is sum(rate(demo_requests_total{outcome="error"}[5m])) / sum(rate(demo_requests_total[5m])), and the 95th percentile latency is histogram_quantile(0.95, sum by (le) (rate(demo_request_duration_seconds_bucket[5m]))). This is also the pattern for model servers: wrap the prediction call in a histogram timer and count outcomes, and you have latency and error dashboards for your model from day one. Many serving frameworks, among them BentoML, KServe and vLLM, expose such metrics already, so check before writing your own.

Short-lived jobs are a special case A batch job that runs for thirty seconds and exits may never be scraped. For those, the client library can push its final values to a small intermediary called the Pushgateway, which Prometheus then scrapes with honor_labels: true. It is meant for batch jobs only, not as a general way to push metrics, because it remembers every series until someone deletes it.
Try it
  1. Install prometheus-client, save the program as app.py and run it.
  2. Add the demo-app job, validate, reload, and confirm the target is UP.
  3. Graph the request rate, the error ratio and the 95th percentile latency using the three queries above.
a request rate of roughly 5 to 50 per second depending on the sleep, an error ratio near 0.05 that wobbles, and a 95th percentile near 0.38 seconds. The numbers match what the code does, which is the moment instrumentation clicks.

How Prometheus stores data, and how long it keeps it

Everything Prometheus collects goes into its own time-series database (TSDB), stored in the directory set by --storage.tsdb.path, which defaults to data/ next to the program. You do not need to manage it day to day, but a simple picture of how it works helps you answer "where did my disk go" and "what happens if it crashes".

New samples first land in an in-memory area called the head block, and to survive a crash each one is also appended to a write-ahead log, the wal directory, which is replayed when the server restarts. About every two hours the head is written out to disk as an immutable block, a directory holding the chunks of samples, an index and a small meta.json. Later, small blocks are merged into larger ones by compaction, which keeps queries efficient. Samples are compressed well: on average a sample takes one to two bytes on disk. That is why a modest server can keep months of history.

Retention

By default Prometheus keeps 15 days of data and then deletes the oldest blocks. You change this in the configuration file, which is the current way since version 3.8:

prometheus.yml
storage:
  tsdb:
    retention:
      time: 30d
      size: 50GB

If you set both, whichever limit is reached first wins. You will still find tutorials that use the flags --storage.tsdb.retention.time and --storage.tsdb.retention.size. They still work but are deprecated, and the file setting takes precedence and can be reloaded without a restart. Expired blocks are cleaned up periodically, so do not expect the disk to shrink the second you lower the number.

To plan disk, use this estimate: the retention in seconds, times the number of samples ingested per second, times about two bytes. You can see the ingest rate with rate(prometheus_tsdb_head_samples_appended_total[5m]). A server scraping 100,000 series every 15 seconds ingests around 6,700 samples per second and needs on the order of a gigabyte a day.

What local storage is, and is not

The local database is fast and simple, but it is a single copy on one disk. It is not replicated, and it is not designed for network file systems such as NFS, which do not give the guarantees it needs. Put it on a local SSD or a block volume, and treat long-term retention and high availability as separate problems with separate answers: remote write to a long-term store, running two identical servers, or both. Those are Senior topics. One final rule for containers: two Prometheus processes must never share a data directory. The second one fails at startup with opening storage failed: lock DB directory: resource temporarily unavailable, because the first holds a lock file.

Try it
  1. List your data directory and identify wal, chunks_head and any block directories with long random-looking names.
  2. Query prometheus_tsdb_head_series and rate(prometheus_tsdb_head_samples_appended_total[5m]).
  3. Estimate how many gigabytes per day your current setup would need, using the formula above.
a few directories in the data folder, a series count in the thousands, and an estimate well under a gigabyte a day. Knowing the number once stops disk surprises later.

Day-to-day operation and the HTTP API

Once a server is running you mostly talk to it through four things: the web interface, promtool, signals, and the HTTP API. The first three you have met. The API matters because everything the interface does, and everything Grafana does, goes through it.

Every query you can type in the box can be sent over HTTP. An instant query uses /api/v1/query, and a range query with a start, end and step uses /api/v1/query_range:

BASH
curl -s 'localhost:9090/api/v1/query' --data-urlencode 'query=up'
BASH
curl -s 'localhost:9090/api/v1/query_range' \
  --data-urlencode 'query=rate(node_cpu_seconds_total{mode="idle"}[5m])' \
  --data-urlencode "start=$(date -d '1 hour ago' +%s)" \
  --data-urlencode "end=$(date +%s)" \
  --data-urlencode 'step=60'

(On macOS, replace the date -d '1 hour ago' part with date -v-1H.) The answer is JSON with a status of success or error, and the data inside it. The same API offers /api/v1/targets for scrape targets, /api/v1/rules and /api/v1/alerts for rules and their state, /api/v1/labels and /api/v1/series for exploring what exists, and /api/v1/status/buildinfo for the version. Wrapping curl around these endpoints is how scripts and other tools integrate with Prometheus. The promtool command can do the same from the terminal:

BASH
./promtool query instant http://localhost:9090 'up'
./promtool check healthy --url=http://localhost:9090

Reading the logs

Logs go to standard error, in a format of time=... level=INFO source=... msg=... since version 3.0, which uses the key=value style of the Go slog package. Older tutorials show a different style starting with ts= and caller=. You can switch to JSON with --log.format=json. To see more detail, the log level is info by default. The command-line flag --log.level still works but has been deprecated since 3.15, and the configuration file now carries the setting, which can be reloaded without restarting:

prometheus.yml
runtime:
  log_level: debug

Watching Prometheus itself

Because Prometheus is scraping itself, you can build a small list of self-checks from day one, and they are the seed of the alerts you will want in production. Whether the last reload worked: prometheus_config_last_reload_successful. How many series are in memory: prometheus_tsdb_head_series. Whether any rule evaluations fail: increase(prometheus_rule_evaluation_failures_total[1h]). And whether notifications reach Alertmanager: increase(prometheus_notifications_errors_total[1h]). A monitoring system that nobody is watching is the classic irony, so sooner or later you will add a rule that alerts when Prometheus itself is unhealthy.

Try it
  1. Run an instant query through curl and read the JSON, finding the metric labels and the value pair.
  2. Fetch /api/v1/targets and find the health field for each target.
  3. Run the same query with promtool query instant and compare.
the same data three ways: a value pair of a Unix timestamp and a string number, targets with health "up", and promtool's terser text. All three read from the same API.

Common errors and how to read them

Most Prometheus problems announce themselves clearly, in the logs, on the Target health page, or in a PromQL error box. The skill is knowing where to look first. This section lists the messages you are most likely to meet as a beginner, with the cause and the fix.

Error loading config followed by a parsing message. The configuration file is not valid YAML or has an unknown field. Run promtool check config prometheus.yml and fix the reported line, most often indentation or a tab character.

scrape timeout greater than scrape interval for scrape config with job name "X". The timeout must not exceed the interval. Lower scrape_timeout or raise the interval.

found multiple scrape configs with job name "X". Two jobs share a name. Names must be unique across the whole configuration.

Target DOWN with connection refused. Nothing is listening at that address. Start the exporter, or check the port. If it runs in a container, check that the target address is reachable from the Prometheus container, because localhost inside a container means the container itself, not your laptop.

Target DOWN with context deadline exceeded. The target did not answer within the scrape timeout. The cause is usually a firewall that silently drops packets, or a target that is too slow to generate its output.

Target DOWN with server returned HTTP status 404. The path is wrong, often because the application serves metrics somewhere other than /metrics. Set metrics_path in the scrape config.

received unsupported Content-Type "text/html" and no fallback_scrape_protocol specified. Since 3.0 Prometheus is strict about the content type of a scrape response. You pointed it at a web page instead of a metrics endpoint, or an exporter sends no proper type. Fix the target, or as a last resort set fallback_scrape_protocol: PrometheusText0.0.4.

listen tcp 0.0.0.0:9090: bind: address already in use. Another program, often a previous Prometheus you forgot to stop, holds the port. Stop it, or start this one with --web.listen-address=:9091.

opening storage failed: lock DB directory. Two servers point at the same data directory. Stop one. With Docker, this also happens when you restart a container while the old one is still shutting down.

opening storage failed: ... permission denied. The user running Prometheus cannot write the data directory. In Docker this appears after switching between image variants, because the default image runs as user nobody and the -distroless image runs as user 65532. Change the ownership of the volume to match.

A PromQL parse error such as 1:1: parse error. The box shows the position. Typical culprits are a missing bracket, a missing quote, or rate(x) without a range, which gives expected type range vector in call to function "rate", got instant vector. Another is vector selector must contain at least one non-empty matcher, which means a selector like {job=~".*"} that could match anything.

query processing would load too many samples into memory in query execution. The query touches more than the limit of 50 million samples. Narrow the selector or the time range, or aggregate earlier.

A graph that is empty though the target is UP. Check that the metric name is spelled the way the exporter spells it, which you can do with curl against the target. Check the label values exactly. Also check that you are not in a time range that predates the data.

After a reload, nothing changed. The new file may have been rejected, in which case the old configuration keeps running. Look at the logs, query prometheus_config_last_reload_successful, and run promtool check config.

Old tutorials will mislead you Prometheus 3.0 removed or changed several things that many pages on the web still show. The flags --enable-feature=agent and --enable-feature=remote-write-receiver are gone, and holt_winters was renamed. The Alertmanager API version 1 is unsupported. Label values of le and quantile are now written in float form, so le="1" is le="1.0". The safest habit is to check the date of a tutorial and confirm its flags against prometheus --help on your own machine.
Try it
  1. Break something on purpose, one at a time: point a job at a port with nothing on it, set a scrape timeout larger than the interval, and type rate(up).
  2. For each one, find where the error shows (logs, Target health, the query box) and read the exact wording.
  3. Fix each and confirm the error disappears.
three different places for three errors, each with a message that names its own cause. Once you have seen an error deliberately, you recognise it instantly when it happens for real.

Putting it all together

Here is one small project that uses everything above, and it is worth building end to end once. Goal: monitor a machine and a small application, record a useful ratio, and alert when the application's error rate climbs or a target stops answering. Allow about an hour.

  1. Make a working folderCreate a directory with the Prometheus files, node_exporter, and the app.py from earlier, plus empty prometheus.yml, rules.yml and alertmanager.yml.
  2. Write the configurationOne file with a 15-second interval, rule files, the Alertmanager address, and three jobs: prometheus, node, demo-app.
  3. Write the rulesTwo recording rules for node utilisation, a recorded error ratio for the app, and two alerts.
  4. Validate everythingRun promtool check config and promtool check rules until both say SUCCESS.
  5. Start and verifyStart the four processes and confirm all three targets are UP.
  6. Break itIncrease the error probability in app.py and watch the alert go pending, then firing, then reach Alertmanager.

The configuration file brings it together:

prometheus.yml
global:
  scrape_interval: 15s
  evaluation_interval: 15s

alerting:
  alertmanagers:
    - static_configs:
        - targets: ["localhost:9093"]

rule_files:
  - "rules.yml"

storage:
  tsdb:
    retention:
      time: 15d

scrape_configs:
  - job_name: "prometheus"
    static_configs:
      - targets: ["localhost:9090"]

  - job_name: "node"
    static_configs:
      - targets: ["localhost:9100"]

  - job_name: "demo-app"
    static_configs:
      - targets: ["localhost:8000"]

And the rules, combining a recording rule for the error ratio with the alerts it feeds:

rules.yml
groups:
  - name: demo
    rules:
      - record: job:demo_errors:ratio_rate5m
        expr: >
          sum(rate(demo_requests_total{outcome="error"}[5m]))
          / sum(rate(demo_requests_total[5m]))

      - alert: DemoHighErrorRatio
        expr: job:demo_errors:ratio_rate5m > 0.10
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "Demo app error ratio is {{ $value | humanizePercentage }}"

      - alert: TargetDown
        expr: up == 0
        for: 2m
        labels:
          severity: page
        annotations:
          summary: "{{ $labels.job }} on {{ $labels.instance }} is down"

Start the pieces in separate terminals, in the order that feels natural: node_exporter, the app, Alertmanager, then Prometheus. Open the Target health page and confirm three UP rows. Graph job:demo_errors:ratio_rate5m. Now change 0.05 in app.py to 0.30, restart it, and follow the alert through its life: pending for two minutes, firing, and then visible in Alertmanager at port 9093. Restore the value and watch it resolve. You have now completed the full loop that production monitoring follows: expose, scrape, store, query, record, alert, notify.

Keep the project Put this folder in a Git repository. Configuration and rules are code: they deserve history, review and a place in CI where promtool check runs on every change. That habit is what separates a monitoring setup that survives staff changes from one that only its creator understands.

What you can now do, and what comes next

You started without knowing what a scrape was. You can now explain the pull model and why it makes a dead service visible. You can read the notation metric{label="value"}, know why a label must have bounded values, and choose between counters, gauges, histograms and summaries. You can install Prometheus on your machine, validate its configuration with promtool, reload it without a restart, and read the Target health page when something is wrong. You can monitor a host with node_exporter, instrument a Python program, and write the PromQL queries for the request rate, error ratio and latency percentile. You can turn a query into a recording rule, write an alert with a sensible for duration, and route it through Alertmanager. You can also decode most of the error messages you will meet in your first month.

Each of those is a real skill that appears in job descriptions for platform, DevOps and MLOps roles across the region, where teams run Prometheus on their own Kubernetes clusters or through managed services from the large cloud providers.

What comes next. The Mid-level guide goes under the surface. It covers service discovery so that targets appear automatically, relabeling to reshape what you collect, the trickier parts of PromQL such as vector matching and subqueries, how to unit-test your rules, how to secure a server, and the integrations with Grafana, Kubernetes and OpenTelemetry. The Senior guide treats Prometheus as a platform: high availability, long-term storage, cardinality control at scale, upgrades, and multi-team governance. Two neighbours are worth reading in parallel. Grafana, the dashboard tool most teams put in front of Prometheus, has its own guide at /student-guides/grafana. And if you run on Kubernetes, the Kubernetes guide is the place to learn what Prometheus will be watching there. For an MLOps engineer, the natural next exercise is to serve a model behind a small API, instrument it with a histogram and a counter as you did here, and build the latency and error alerts for it.

Sources