Skip to content
Back to student guides
GrafanaDevOpsObservability3 levels126 sectionsCovers Grafana 13.2

The Complete Grafana Guide

Visualise metrics, logs and traces in Grafana dashboards and alerts. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
21sections
32examples

This is part one of three. It covers everything you need to do real work with Grafana, not a teaser. By the end you can install it, connect it to a data source, explore data with queries, build a dashboard with variables, set up an alert that reaches a human, keep your configuration in files instead of clicks, and call the API from a script. Mid-level and Senior take the same topics further; nothing here is thrown away.

The guide is written against Grafana 13.2 (the latest patch at the time of writing is 13.2.3, published on 29 September 2026). Grafana moves fast: a new minor version arrives every six to ten weeks and patches roughly every two weeks, so a few commands that older tutorials show no longer work. Wherever that matters, this guide says so and shows the current form.

Each section ends with a Try it task. Do them as you go. Grafana is a visual tool, and the concepts only stick once you have watched your own panel draw a line, your own alert turn red, and your own dashboard survive a restart.

What Grafana is, and the problem it solves

Grafana is a web application that asks your data stores questions and draws the answers. You point it at a place where numbers, logs, or traces already live (Prometheus, a PostgreSQL database, Loki, Elasticsearch, a cloud monitoring service) and it gives you charts, tables, dashboards, and alerts on top.

The most important sentence in this guide is this one: Grafana is not a datastore. It does not collect your metrics and it does not hold your logs. The data stays in the system that produced or stored it. Grafana only keeps its own bookkeeping: your dashboards, users, data source settings, alert rules, and preferences. That bookkeeping lives in a small database that Grafana manages for itself.

YOUR SYSTEMSapps, servers, models
→
A DATA STOREPrometheus, Loki, SQL
→
GRAFANAqueries and draws
→
A HUMANdashboard or alert

To see why this design exists, think about what monitoring looked like before. A typical team had a metrics tool with its own graph page, a log tool with its own search page, a database console for business numbers, and a cloud console for the infrastructure bill. During an incident, an engineer had four browser tabs, four query languages, and four different time pickers, and had to line up timestamps by eye. Each tool drew its own charts, and each had its own idea of users and permissions.

Grafana's answer is to separate storing from looking. The stores stay specialised (Prometheus is very good at metrics, Loki at logs), and one interface sits on top that speaks to all of them. One time picker covers every panel. One dashboard can show a latency graph from Prometheus next to the error logs from Loki next to a business count from PostgreSQL. One alerting system watches all of them.

For an MLOps engineer this matters in concrete ways. A model-serving endpoint produces request latency and error counts (metrics), prediction logs (logs), and request traces. A training pipeline produces GPU utilisation and loss curves. A feature pipeline produces freshness and row-count numbers in a warehouse. None of that belongs in one database, but all of it belongs on one screen when something breaks at two in the morning. Teams in the Gulf and Egypt who run these systems in a regional cloud or an on-premises cluster for data-residency reasons can also run Grafana in the same place, since it is ordinary software you host yourself; nothing about your data needs to leave your network.

What people use Grafana for:

📈

Dashboards

One screen that shows whether a service, a pipeline, or a model is healthy right now.

🔎

Investigation

Ad hoc queries in Explore while you chase a problem, without building anything permanent.

🔔

Alerting

Rules that watch your data and notify Slack, email, or a pager when something crosses a line.

🧩

One interface

Metrics, logs, traces, SQL, and cloud data side by side, with a single time range.

Grafana comes in three forms, and it helps to know the names now. Grafana OSS is the open-source edition (AGPLv3 licence) that you host. Grafana Enterprise is the same software plus licensed extras such as fine-grained role-based access control, SAML single sign-on, and some extra data sources; the Enterprise image runs with all the open-source features and needs no licence for that. Grafana Cloud is the hosted service run by Grafana Labs. This guide works on the first two and notes where a feature is Enterprise or Cloud only.

Try it
  1. Pick one system you know. Write down three questions you would ask about it during an outage (for example, "is latency up", "are errors up", "did a deploy just happen").
  2. For each question, write where the answer lives: a metrics store, a log store, a database, or a cloud console.
the answers live in more than one place. That spread is exactly what Grafana exists to bring onto one screen.

The mental model: five nouns

Grafana has many features, but almost everything reduces to five nouns. Learn these and the interface stops being a maze.

A data source is a saved connection to a backend: a type (Prometheus, PostgreSQL, Loki), an address, and credentials. Every data source has a UID, a short string identifier. You will see UIDs in URLs, in provisioning files, and in the API. Older tutorials use a numeric ID instead; that form is legacy, and since Grafana 13 the numeric-ID data source APIs are disabled by default, so always work with UIDs.

A query is the question you ask a data source. The language depends on the data source: PromQL for Prometheus, LogQL for Loki, SQL for PostgreSQL. Grafana sends the query, and the data source answers with data frames: tables with typed columns. A time series is just a frame with a time column and one or more value columns. You rarely need to think about frames until you use transformations, but knowing that every answer is a table explains much of Grafana's behaviour.

A panel is one rendered visualisation: a time series line chart, a single big number (called a Stat), a gauge, a table, a log viewer, a heatmap. A panel is made of queries (what to ask), optional transformations (how to reshape the answer), and display options (how to draw it).

A dashboard is a saved page of panels, together with a time range, variables, and links. A dashboard is stored as JSON and identified by a UID. Dashboards live inside folders, and folders are where you assign who may see or edit them.

An alert rule is a saved query plus a condition, evaluated on a schedule. When the condition holds, the rule produces an alert, and Grafana routes that alert to a contact point (Slack, email, a webhook) according to a notification policy.

Data sourceconnection + UID
is queried by
QueryPromQL, SQL, LogQL
feeds
Panelone visualisation
lives in
Dashboardpanels + variables

A few supporting nouns come up constantly. A plugin extends Grafana: a data source plugin adds a new backend type, a panel plugin adds a new visualisation, and an app plugin bundles pages and features. An organisation (org) is a hard boundary inside one Grafana instance: data sources, dashboards, and teams belong to exactly one org. A fresh install has one org called "Main Org." and most people never need a second. A team is a group of users that you grant permissions to in one step. Explore is the scratch pad where you run queries without building a panel. A variable is a dropdown at the top of a dashboard that changes the queries underneath it.

Finally, know where Grafana keeps its own state. By default it uses SQLite, a single file inside the data directory. That is fine for learning and for a single-user setup. For anything shared and important, Grafana supports MySQL 8.0 or later and PostgreSQL 12 or later. You never put your monitoring data in this database; it holds only Grafana's own configuration.

Read this twice. If you delete Grafana's database, you lose dashboards, users, and alert rules, but your metrics and logs are untouched because they live elsewhere. If you delete Prometheus's storage, you lose your metrics, but your dashboards still exist and simply show "No data". Keeping these two failure modes separate in your head prevents a lot of panic.
Try it
  1. Without looking back, write the five nouns (data source, query, panel, dashboard, alert rule) on paper.
  2. Next to each, write one sentence that describes it and one example from a system you know.
each sentence fits on a line. If one does not, reread its paragraph; the rest of the guide leans on these five.

Versions, editions, and what changed recently

Before installing anything, it is worth five minutes on versions, because Grafana's recent history contains a few changes that break older instructions. You will meet blog posts and videos from 2022 or 2023 that describe a different product than the one you are about to install.

The numbering is simple: Grafana 12 was released in 2025, and Grafana 13 arrived in April 2026, followed by 13.1 in July and 13.2 on 18 August 2026. Several maintenance branches receive patches at the same time: as of late September 2026 those are 13.2, 13.1, 13.0, and the last 12.x line, 12.4. Patch releases are frequent and often security-motivated; the 13.2.3 patch contains only security fixes. The practical rule is to pin an exact version tag such as grafana/grafana-enterprise:13.2.3 in anything you keep, and never use a floating tag like latest outside an experiment, because latest will silently change underneath you on the next pull.

The changes from Grafana 13.0 that matter to a beginner:

  • One binary, two subcommands. The old stand-alone commands grafana-server and grafana-cli were removed. You now run grafana server and grafana cli. The Linux service is still named grafana-server (the systemd unit name did not change), which confuses everyone at first: the unit keeps its name, the removed thing is the standalone command.
  • The Image Renderer plugin is gone. Turning dashboards into PNG images (for scheduled reports and alert screenshots) now needs the separate remote rendering service. You will not need this as a beginner, but old guides that say "install the renderer plugin" are obsolete.
  • /api is deprecated in favour of /apis. Grafana is moving to Kubernetes-style resource APIs. The old /api routes still work and are not being switched off, but they no longer receive updates. This guide uses the familiar /api endpoints for a beginner's needs and shows where the newer form appears.
  • Dynamic dashboards and Git Sync reached general availability. Dynamic dashboards add rows, tabs, and automatic layouts. Git Sync lets a Git repository and a Grafana folder stay in step. Both come up again near the end.
  • Older removals still bite. Angular-based plugins no longer load (removed in Grafana 12), legacy alerting is gone (removed in 11), and API keys were replaced by service accounts.
Trap: an old tutorial. If a tutorial tells you to run grafana-cli plugins install, create an "API key", or use the "Graph" panel, it predates these changes. Stop and look for its modern equivalent: grafana cli plugins install, a service account token, and the Time series panel.

There is also a warning about upgrades worth memorising early. Grafana 13.0.0 had a migration bug that could lose or revert dashboards and folders when upgrading from 12.x with Git Sync enabled. If you ever upgrade across a major version, back up the database first, use the latest patch of the target line rather than the first release, and try it in a throwaway copy before the real one.

Try it
  1. Open the "What's new" page for Grafana 13 from the Sources list at the end of this guide.
  2. Find the breaking-changes section and read the list once.
  3. Write down which items could affect a tutorial you followed in the past.
at least one item (a renamed command, a removed plugin, or a changed API) that explains an error you once hit.

Installing Grafana and checking the setup

There are several ways to install Grafana. They all produce the same program. Pick the one that matches where you will run it, and do not agonise over it: you can move your dashboards later.

Grafana needs little: a minimum of 512 MB of memory and one CPU core for a single-user setup. The sizing guidance for a shared instance is roughly 2 cores and 2 to 4 GB for fewer than 25 concurrent users, rising from there. Supported platforms are Debian and Ubuntu, RHEL and Fedora, SUSE, macOS, and Windows.

Docker (the quickest way to start)

If you have Docker (see the Docker guide), this is the fastest route and leaves nothing behind on your machine.

BASH
docker volume create grafana-storage
docker run -d -p 3000:3000 --name=grafana \
  --volume grafana-storage:/var/lib/grafana \
  grafana/grafana-enterprise:13.2.3

Read it line by line. The first command creates a named volume, a piece of storage Docker manages, so that Grafana's database survives when you delete and recreate the container. Without it, removing the container removes your dashboards. -d runs the container in the background. -p 3000:3000 publishes the container's port 3000 on your machine's port 3000, which is Grafana's default. --volume grafana-storage:/var/lib/grafana attaches the volume at the directory where Grafana keeps its data. The image grafana/grafana-enterprise holds every open-source feature and needs no licence for them; if you prefer the pure open-source build, the image grafana/grafana is published too. Note the pinned tag at the end.

You can change settings at start-up with environment variables, which are explained in the configuration section. For example, -e GF_LOG_LEVEL=debug raises the log detail, and -e GF_PLUGINS_PREINSTALL=grafana-clock-panel@1.0.1 installs a plugin at start.

Docker Compose (better for a real project)

When Grafana will sit beside other services such as Prometheus, a Compose file is cleaner than a long docker run line.

docker-compose.yml
services:
  grafana:
    image: grafana/grafana-enterprise:13.2.3
    container_name: grafana
    restart: unless-stopped
    ports:
      - "3000:3000"
    volumes:
      - grafana_storage:/var/lib/grafana

volumes:
  grafana_storage: {}

Start it with docker compose up -d. We extend this file in the first-project section.

Linux packages

On Debian or Ubuntu, add Grafana's package repository and install from it:

BASH
sudo apt-get install -y apt-transport-https wget gnupg
sudo mkdir -p /etc/apt/keyrings
sudo wget -O /etc/apt/keyrings/grafana.asc https://apt.grafana.com/gpg-full.key
sudo chmod 644 /etc/apt/keyrings/grafana.asc
echo "deb [signed-by=/etc/apt/keyrings/grafana.asc] https://apt.grafana.com stable main" | sudo tee -a /etc/apt/sources.list.d/grafana.list
sudo apt-get update
sudo apt-get install grafana-enterprise
sudo systemctl daemon-reload
sudo systemctl enable --now grafana-server

The commands fetch the repository's signing key so the package manager can verify what it installs, register the repository, then install the package and start the service. Install grafana instead of grafana-enterprise if you want the pure open-source package. On RHEL, Fedora, Rocky, or Alma you create /etc/yum.repos.d/grafana.repo pointing at https://rpm.grafana.com, import the key, and run sudo dnf install grafana-enterprise; the exact file is on the official install page listed in Sources. The same systemctl enable --now grafana-server starts it.

macOS and Windows

On macOS, Homebrew is the easiest route:

BASH
brew update && brew install grafana
brew services start grafana

Homebrew's formula may lag behind the newest patch, so check what you got with grafana --version. On Apple Silicon the configuration file lives at /opt/homebrew/etc/grafana/grafana.ini. On Windows, download the MSI installer or the ZIP from the download page; with the ZIP, right-click it, open Properties, and choose Unblock before extracting. From Grafana 13 the command to start the server from the extracted folder is bin\grafana.exe server.

Shortcut. If you only want to learn, use Docker. There is nothing to uninstall: docker rm -f grafana removes the container and docker volume rm grafana-storage removes the data.

Checking that it works

Whichever route you took, verify in three ways. First, ask the health endpoint, which needs no login:

BASH
curl -s http://localhost:3000/api/health
JSON
{
  "database": "ok",
  "version": "13.2.3",
  "commit": "…"
}

"database": "ok" means Grafana reached its own database. If it says failing, Grafana is running but its storage is broken (a common cause is a permissions problem on the data directory). Second, open http://localhost:3000 in a browser. Third, read the logs: docker logs grafana for a container, journalctl -u grafana-server for a Linux service. You are looking for a line saying the HTTP server is listening on port 3000.

The first login is admin with password admin. Grafana immediately asks you to choose a new password; do it, because that default is the first thing anyone scanning for Grafana instances tries.

Trap: the password that "does not change". The admin password from settings only applies the very first time Grafana creates its database. After that, the password lives in the database, and editing GF_SECURITY_ADMIN_PASSWORD changes nothing. To reset a forgotten password, use grafana cli admin reset-admin-password 'NewPassword' on the host (for a container, run it with docker exec).
Try it
  1. Start Grafana with Docker using a pinned tag and a named volume.
  2. Run the curl health check and confirm "database": "ok".
  3. Log in, change the password, then run docker rm -f grafana and start the same command again.
your new password still works, because the volume kept the database while the container was replaced.

A guided tour of the interface

Once you are logged in, spend ten minutes looking around before you build anything. Grafana's interface is organised around a left-hand navigation menu, and knowing what each entry is for saves hours of hunting. The exact wording of menu items shifts slightly between releases (the homepage was redesigned in 13.2, for instance), so if a label differs, search by purpose.

  • Home is the landing page. It shows starred and recent dashboards.
  • Dashboards lists folders and dashboards, with search, tags, and a button to create new ones or import existing ones.
  • Explore is the ad hoc query page. You pick a data source, type a query, and see results immediately, without saving anything.
  • Alerting holds alert rules, contact points, notification policies, silences, and mute timings.
  • Connections is where data sources live (Connections, then Data sources), and where you can add new ones.
  • Administration holds users, teams, service accounts, plugins, settings, and the stats page. Only admins see most of it.

The command palette opens with Ctrl+K (Cmd+K on a Mac) since Grafana 13.2. Type a dashboard name, a page, or an action and jump straight to it. It is the fastest way to move around once you know what you want.

Two controls appear on almost every dashboard and deserve an early introduction. The time range picker at the top right sets the window all panels show: "Last 6 hours", "Last 7 days", or an absolute range. The refresh picker beside it tells the dashboard to re-run its queries every so often (5 seconds, 1 minute, and so on) or never. Both apply to the whole dashboard, which is what makes a dashboard feel like one instrument rather than a pile of charts.

Finally, look at your profile (your avatar) and Administration, then General, then Stats and licensing. The second page shows the exact version of the server you are talking to, which is worth checking whenever behaviour does not match the documentation you are reading.

Try it
  1. Open each of the six menu areas above and write one line about what it shows.
  2. Press Ctrl+K (or Cmd+K) and jump to Explore using only the keyboard.
  3. Open Administration, then General, then Stats and licensing, and note the version.
the version number matches the tag you pinned, and you can navigate without touching the mouse.

Your first data source: Prometheus

An empty Grafana draws nothing, so the first real step is to give it something to ask. We will use Prometheus, the most common metrics store in the Grafana world (see the Prometheus guide for the full story). Prometheus regularly pulls numbers from your services, stores them as time series, and answers queries in a language called PromQL. Prometheus can scrape Grafana itself, which gives us real data with nothing else to install.

Extend the Compose file so that Prometheus runs next to Grafana and scrapes both. First the Prometheus configuration:

prometheus.yml
global:
  scrape_interval: 15s

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ["prometheus:9090"]

  - job_name: grafana
    metrics_path: /metrics
    static_configs:
      - targets: ["grafana:3000"]

A scrape is one HTTP request in which Prometheus fetches the current numbers from a target. scrape_interval: 15s means it does that every fifteen seconds. The targets use service names (prometheus, grafana), not localhost, because inside a Compose network every container has its own idea of localhost. Keep that in mind; it is the most common first error.

Now the Compose file:

docker-compose.yml
services:
  grafana:
    image: grafana/grafana-enterprise:13.2.3
    container_name: grafana
    restart: unless-stopped
    ports:
      - "3000:3000"
    environment:
      GF_METRICS_ENABLED: "true"
    volumes:
      - grafana_storage:/var/lib/grafana

  prometheus:
    image: prom/prometheus
    container_name: prometheus
    restart: unless-stopped
    ports:
      - "9090:9090"
    volumes:
      - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro
      - prometheus_data:/prometheus

volumes:
  grafana_storage: {}
  prometheus_data: {}

GF_METRICS_ENABLED turns on Grafana's own /metrics endpoint, which is how Prometheus will read Grafana's numbers. (This environment variable naming rule is explained in the configuration section.) In a real project, pin the Prometheus image to an exact tag too. Bring it up with docker compose up -d, then open http://localhost:9090/targets and confirm both targets say UP.

Next, tell Grafana about Prometheus. In the UI choose Connections, then Data sources, then Add data source, pick Prometheus, and fill in one field that matters:

  • URL: http://prometheus:9090

Leave the rest at its defaults and press Save and test. A green banner saying the data source is working means Grafana reached Prometheus and got a sensible answer.

Trap: localhost is the wrong address. The data source URL is opened by the Grafana server, not by your browser. Inside a container, http://localhost:9090 points at Grafana's own container, where nothing listens on 9090, and you get "connection refused". Use the service name from your Compose file. The same rule applies to a database or any other backend: what matters is what the Grafana process can reach.

The access mode on a data source form used to offer "Server" and "Browser". Server mode, where Grafana's backend makes the request on your behalf, is the default and the only sensible choice; it keeps credentials out of the browser and lets Grafana apply permissions and timeouts.

If the test fails, read the message. "connection refused" means nothing listens at that address from Grafana's point of view. "no such host" means the name does not resolve (a typo, or the containers are on different networks). "401" means the backend demands authentication you have not configured. A slow test that ends in a timeout usually means a firewall between the two.

Try it
  1. Create the two files above and run docker compose up -d.
  2. Confirm both targets show UP at http://localhost:9090/targets.
  3. Add the Prometheus data source with the URL http://prometheus:9090 and press Save and test.
  4. Edit the URL to http://localhost:9090 and test again, then fix it back.
a green success banner for the right URL and a connection-refused error for the wrong one, which is the error you will meet most often in real life.

Explore: asking questions without building anything

Before you design any dashboard, use Explore. It is the place to work out what to ask, and it is where experienced people spend much of their time, because a dashboard is only the frozen answer to a question someone already figured out.

Open Explore, choose your Prometheus data source, and type the simplest possible PromQL query:

TEXT
up

up is a metric Prometheus creates for every target it scrapes: 1 if the last scrape succeeded, 0 if it failed. Press Run query (or Shift+Enter). You see a line for each target, sitting at 1. Each line is a time series: a metric name plus a set of labels (key and value pairs such as job="grafana") that identify it, and a stream of timestamped values.

Labels are the heart of PromQL. To select only one job, filter by label:

TEXT
up{job="grafana"}

Now something more useful. Prometheus stores many metrics as counters: numbers that only ever go up, like the total number of requests served. A raw counter is rarely interesting; what you want is how fast it grows. That is what rate does:

TEXT
rate(prometheus_http_requests_total[5m])

Read it inside out. prometheus_http_requests_total is a counter of requests Prometheus has served. [5m] selects the last five minutes of samples for each series. rate(...) turns that into the average per-second increase over that window. The result is requests per second, per handler.

To collapse the series into one total, aggregate:

TEXT
sum(rate(prometheus_http_requests_total[5m]))

And to keep a breakdown by one label while summing away the rest:

TEXT
sum by (handler) (rate(prometheus_http_requests_total[5m]))

Explore has a Builder mode that lets you pick the metric, labels, and functions from dropdowns, and a Code mode where you type PromQL directly. Start in Builder to learn the vocabulary; switch to Code when you know what you want. The toggle sits at the top right of the query editor. Explore also keeps a query history, so a query you ran yesterday is one click away, and it has a split view for comparing two queries (or two data sources) side by side.

There is one more thing worth understanding early, because it confuses beginners: the step. A chart has a limited number of pixels, so Grafana asks Prometheus for a limited number of points across your time range. When you zoom out to seven days, each point covers a much longer period than when you look at the last hour. Grafana provides built-in variables for this, and the most important one is $__rate_interval, which picks a safe window for rate based on the step and the scrape interval. In a dashboard panel you should write:

TEXT
sum by (handler) (rate(prometheus_http_requests_total[$__rate_interval]))

instead of hard-coding [5m]. A fixed window that is too small for the step produces gaps ("No data" holes) when you zoom out; $__rate_interval adapts.

Shortcut. Press the "Metrics" drilldown or the metrics browser in Explore to search what metrics exist when you do not know the name. In a fresh Prometheus, typing a few letters in the Builder's metric dropdown autocompletes from what the server really has.
Try it
  1. In Explore, run up, then up{job="grafana"}.
  2. Run rate(prometheus_http_requests_total[5m]), then wrap it in sum by (handler).
  3. Reload the Prometheus targets page a few times, then re-run and watch the rate change.
the request rate for Prometheus's own pages rises after you generate traffic. You have just watched a counter turn into a rate.

Building your first dashboard

An investigation in Explore becomes durable once it lives in a dashboard. You will now build one with three panels.

Create a dashboard from Dashboards, then New, then New dashboard, and choose to add a visualisation. Grafana asks which data source to use: pick Prometheus. You land in the panel editor, which has three areas: the preview on top, the query editor below, and the options pane on the right. Everything you do in it falls into those three areas.

Panel one: a time series of request rate. Paste the query from the last section into the query editor:

TEXT
sum by (handler) (rate(prometheus_http_requests_total[$__rate_interval]))

The preview draws one line per handler. On the right, find the Title field and call it "Prometheus requests per second". Under Standard options, set the Unit to "requests/sec (rps)" so the axis reads properly. Good units are the cheapest upgrade you can give a dashboard: a number with no unit forces everyone to guess.

Panel two: a Stat for target health. Add a new visualisation. Change the visualisation type (top right of the options pane) to Stat, and use this query:

TEXT
count(up == 1)

up == 1 keeps only the healthy targets and count counts them, so the Stat shows a single big number. Stat panels are the right choice for "what is it right now". Under Thresholds you can make the value go green above a number and red below it; a colour tells a viewer more quickly than a digit does.

Panel three: a gauge or a table. Add a third panel. Try a Table with the query up, and set Format to Table in the query options so the result arrives as rows rather than as a time series. Tables are the right choice when the question is "which one" rather than "how it changes".

Choosing a visualisation comes down to the question you are answering:

Match the chart to the question

  • "How does it change over time?" is a Time series
  • "What is it right now?" is a Stat or a Gauge
  • "Which ones are unhealthy?" is a Table
  • "What did it log?" is a Logs panel

Common mismatches

  • A Time series for a single current value
  • A pie chart with fifteen slices
  • Twenty unlabelled lines on one graph
  • No units and no title on any panel

When you finish, press Save dashboard, give it a name such as "Monitoring stack overview", and choose a folder. The dashboard is now stored in Grafana's database. Two habits are worth forming immediately: give every panel a title and a unit, and put the most important panel top-left, because that is where eyes land first.

A word on layout. Since Grafana 13, dynamic dashboards are generally available, and every dashboard is rendered by the newer Scenes engine. In practice that means you can group panels into rows (collapsible sections) or tabs, and you can select several panels and group them in one step (a 13.2 feature). Existing dashboards are migrated automatically when opened. One caution: once a dashboard adopts the newer dynamic schema, it cannot be downgraded to an older Grafana. That is fine if you only run 13.x, and worth remembering if you share dashboards with colleagues on older servers.

Dashboards are JSON underneath. Open the dashboard's settings and find JSON Model to see the whole thing: panel positions, queries, options, and variables. You do not need to read it fluently now, but two facts make it valuable. Everything you can click, you can also put in a file (which is how provisioning and Git work later), and when something looks wrong, the JSON shows exactly what Grafana stored.

Try it
  1. Build the three panels above. Give each a title and choose sensible units.
  2. Set the Stat's thresholds so it is green at 2 and red below 2.
  3. Save the dashboard into a folder, then open Settings and read the JSON Model.
a saved dashboard with three panels, and a JSON document you can search for the exact query text you typed.

Panels in depth: options, overrides, and transformations

A panel looks simple, but three layers of control sit inside it, and confusing them is a classic beginner stumble. Understanding the layers also tells you where to look when a chart does not look right.

Layer one is the query. It decides what data arrives. If the data is wrong (missing series, wrong numbers), fix it here. The query editor also has query options: the maximum data points, the minimum interval, the relative time range, and the format (time series, table, or heatmap). You can add several queries to one panel, labelled A, B, C, and each draws its own series.

Layer two is transformations. A transformation reshapes the data after it arrives, on the Grafana side, before it is drawn. The Transform tab in the panel editor lists them. A few are worth knowing on day one:

  • Organize fields renames, reorders, and hides columns. It is the tool you reach for when a table has ugly column names.
  • Filter data by values drops rows that do not match a condition.
  • Group by aggregates rows, like a SQL GROUP BY.
  • Join by field merges results from two queries on a shared column, for example a host name.
  • Add field from calculation creates a new column from existing ones, such as a ratio.
  • Reduce collapses each series to a few statistics (last, max, mean), which is how you build a compact table of time series.

Transformations run in a pipeline. Each one receives the output of the previous one, not the original data, so order matters: hiding a column before you calculate from it breaks the calculation. When a transformed result surprises you, disable transformations one by one from the bottom up and watch where the data changes.

Layer three is the field configuration: the Standard options (unit, decimals, min, max, display name), thresholds, value mappings (turn 0 and 1 into "down" and "up"), and overrides. An override applies a setting to some fields only: "for the series named errors, use red and draw as bars". Overrides are how you make one chart carry two different units or styles without splitting it.

Where does the fix belong? Wrong numbers are a query problem. Wrong shape (columns, joins, names) is a transformation problem. Wrong look (colour, unit, line style) is a field configuration problem. Asking "which layer?" before clicking around shortens debugging a lot.

The other decision is when to do work in the query and when to do it in a transformation. If the data source can do it (PromQL aggregation, SQL GROUP BY), do it there: it moves less data across the network and scales better. Use transformations when you need to combine results from different data sources, or when the source cannot express what you need.

A few more panel features pay off quickly. Legends can show values (last, max, mean) in a table beside the graph, which turns a chart into a readable summary. Tooltips can show all series at once or only the one under the cursor. Panel links and data links let a click jump to another dashboard, carrying the time range and variable values with it. Library panels let you define a panel once and reuse it in many dashboards, and editing the library panel updates all of them. Annotations draw vertical markers on time series: a deployment, an incident, a retraining run. Adding a marker by Ctrl-clicking a graph is the quickest way to start, and annotations can also come from a query, so every deploy in your database shows up on every latency graph.

The visualisations you will use most as a beginner, and when:

📉

Time series

The workhorse. Lines, bars, or points over time. Replaces the old Graph panel.

🔢

Stat and Gauge

One current value with colour. The gauge was revamped and is generally available in 13.

📋

Table

Rows and columns, good for "which host" questions and for debugging data shape.

📜

Logs

A scrollable list of log lines from Loki, Elasticsearch, or another log source.

Try it
  1. On your table panel, open Transform and add Organize fields. Hide every column except job and the value.
  2. Rename the value column to "Healthy".
  3. Add a value mapping so 1 shows as "up" (green) and 0 as "down" (red).
a clean two-column table. The data never changed; you only reshaped and restyled it, which is the point of the three layers.

Time ranges, refresh, and variables

A dashboard that shows one fixed service is useful once. A dashboard that lets a viewer choose the service, the environment, or the model is useful for a whole team, and that is what variables do. Combined with the time range, they turn a dashboard into a reusable tool.

The time range

Every panel shares the dashboard's time range. A few points save confusion. Relative ranges like "Last 6 hours" move with the clock; absolute ranges are fixed and good for reproducing an incident. Dragging across a graph zooms the whole dashboard into that selection. The time zone can be browser local, UTC, or a named zone, and UTC is the safest for teams spread across countries, because everyone reads the same timestamps. Refresh intervals have a floor: the server setting min_refresh_interval (5 seconds by default) stops people from hammering a backend with very fast refresh.

Variables

A variable is defined in the dashboard settings and shows as a dropdown at the top. Its value is substituted into queries wherever you write $name (or ${name}). The kinds you will use first:

  • Query variables fill their options by asking a data source, for example "all values of the job label".
  • Custom variables have a fixed list you type in, such as dev,staging,prod.
  • Interval variables offer step sizes like 1m,5m,1h.
  • Data source variables let the viewer switch between data sources of one type.
  • Text box and constant variables hold a free-form or fixed value.

Let us add a job variable to the dashboard. In Settings, open Variables and add a Query variable named job, with the data source set to Prometheus and this query:

TEXT
label_values(up, job)

label_values(up, job) returns every distinct value of the job label on the up metric, which here is prometheus and grafana. Turn on Multi-value and Include All option so viewers can pick several or everything. Now use it in the first panel's query:

TEXT
sum by (handler) (rate(prometheus_http_requests_total{job=~"$job"}[$__rate_interval]))

Notice =~ instead of =. It is the regular-expression match, and it is required because a multi-value variable expands into a pattern like (prometheus|grafana). Using = with a multi-value variable is a classic error: the query returns nothing, because no label value literally equals the whole pattern.

Variables can depend on each other. A second variable can filter by the first, for example the instances that belong to the chosen job, so the choices always make sense together. The order of variables in the list matters because one can only reference variables above it.

Interpolation has formats. ${job:csv} renders a comma-separated list, ${job:pipe} a pipe-separated list, and ${job:regex} a regex-safe pattern. You can ignore these until a data source expects something other than the default; when a multi-value selection produces a syntax error in SQL, a format such as ${var:singlequote} is usually what fixes it. Grafana also provides built-in variables you did not define: $__interval, $__rate_interval, $__range, $__from, $__to, $__dashboard, $__org, and $__user.

Trap: a variable that matches nothing. If a panel says "No data" after you add a variable, check three things in order: is the variable's own preview populated (Settings, Variables, the preview table)? Did you use =~ for a multi-value variable? Did the value you picked actually exist in the selected time range? Variables that query a data source only offer values seen within the time range.

Newer Grafana versions add more ways to filter: group-by and filter controls that let viewers slice by label without writing variables, multi-property variables, and variables scoped to a single row or tab. They are useful, but the classic dashboard variable above is the thing every Grafana engineer knows, so learn it first.

Try it
  1. Add the job query variable with Multi-value and Include All enabled.
  2. Update your request-rate panel to filter with job=~"$job".
  3. Select only grafana, then change =~ to = with both selected, and observe the result.
the panel narrows to one job when you pick one, and goes empty when you combine a multi-value selection with a plain equals sign.

A second kind of data: SQL and logs

Grafana is not only for metrics. Seeing one non-metrics source early stops you from thinking of it as "a Prometheus viewer". The two to know about are SQL databases and Loki for logs.

PostgreSQL and MySQL

For a SQL database, add a data source of type PostgreSQL (or MySQL), and fill in the host and port, the database name, a user, and a password. Create a dedicated read-only database user for Grafana. Anyone who can edit a dashboard can write any query against that connection, so a user that can only SELECT from the tables you chose limits the damage of a mistake or a malicious edit.

A SQL panel query returns rows. For a time series, the query must return a column named time (or use a time macro) and one or more numeric columns:

SQL
SELECT
  $__timeGroupAlias(created_at, '1h'),
  count(*) AS predictions
FROM prediction_log
WHERE $__timeFilter(created_at)
GROUP BY 1
ORDER BY 1

The two macros do real work. $__timeFilter(created_at) expands into a condition that restricts rows to the dashboard's current time range, so the query gets cheaper when you zoom in. $__timeGroupAlias(created_at, '1h') buckets timestamps into hours and names the column time. Without them, you would hard-code dates and the panel would not respond to the time picker. This is how an MLOps team can chart "predictions served per hour" or "rows ingested per day" straight from a warehouse table.

Loki and logs

Loki is Grafana Labs' log store. You query it with LogQL, which looks like PromQL with a log-filtering step. A query has a stream selector in curly braces and optional filters:

TEXT
{job="api"} |= "error"

{job="api"} selects the stream of logs with that label, and |= "error" keeps only lines containing the text. LogQL can also turn logs into numbers: rate({job="api"} |= "error" [5m]) gives errors per second, which you can chart or alert on. Put a logs panel beside a metrics panel on the same dashboard and you can see a spike and read its cause in one glance. The ELK stack guide covers the main alternative for log search, and Grafana can query Elasticsearch directly as well.

The general lesson is that once a data source is configured, the rest of Grafana (panels, variables, alerting, permissions) behaves the same way. You learn one tool and apply it to many stores. There are dozens of data source plugins: CloudWatch, Azure Monitor, Google Cloud Monitoring, InfluxDB, Elasticsearch and OpenSearch, and more.

Shortcut. To chart what an ML service is doing without a metrics pipeline, point Grafana at the database that already stores predictions or evaluation results, with a read-only user, and start with a count-per-hour query like the one above.
Try it
  1. Run a PostgreSQL container, create a table with a created_at timestamp column, and insert a few dozen rows.
  2. Add a PostgreSQL data source using a read-only user.
  3. Paste the SQL above (adapting the table name) into a new time series panel.
a bar or line of rows per hour that changes when you zoom the time picker, proving the macros are wired to the dashboard's time range.

Folders, users, and permissions

The moment more than one person uses Grafana, the question becomes who can see and change what. Grafana's model has three layers, and a beginner only needs to know them at a high level.

Organisation roles are the coarse layer. Every user has a role in each organisation they belong to: Viewer can look at dashboards and run queries, Editor can also create and change dashboards and alert rules, and Admin can also manage data sources, users, and settings for the organisation. There is also a None role with no access by default. Separate from these, a Grafana server admin flag marks someone who manages the whole instance (all organisations); it is distinct from an organisation Admin, which surprises people.

Folder and dashboard permissions are the fine layer. Folders are the main permission boundary: you can give a team View, Edit, or Admin on a folder, and everything inside inherits it. Organising dashboards into folders per team or per system, then granting permissions per folder, is far easier to maintain than granting per dashboard. Folders can nest.

Teams group users so you grant permissions once to "ml-platform" instead of to each person. Role-based access control (RBAC), with custom roles and data source permissions, is an Enterprise and Cloud feature; in the open-source edition you combine organisation roles with folder permissions.

Two security facts every beginner should hold on to:

Trap: Viewers can query anything. A Viewer is restricted in what they can change, not in what they can ask. In the open-source edition, anybody in an organisation can run arbitrary queries against any data source in that organisation, even if a dashboard hides the panel. Treat data source credentials as visible to your whole user base: give Grafana read-only credentials scoped to what you are happy for every user to query.
Trap: anonymous access. The [auth.anonymous] setting lets people use Grafana without logging in. It is off by default, and for good reason: it exposes every data source in that organisation to anyone who can reach the URL. Only enable it on a network you fully control, and never on the public internet.

How users get in matters too. Self-sign-up is off by default ([users] allow_sign_up = false), so an admin creates users or connects a login provider. Grafana supports logging in through GitHub, GitLab, Google, Azure AD, Okta, a generic OAuth or OIDC provider, LDAP, and others; SAML and automatic team sync are Enterprise features. Grafana has no built-in multi-factor authentication, so if you need it, use an identity provider that provides it. A small team can start with local users; anything exposed beyond a single-user lab should use single sign-on.

For automation you do not use a person's account. You create a service account, a non-human identity with an organisation role, and give it tokens. That topic comes up in the API section.

Try it
  1. Create two folders, "Team A" and "Team B", and move your dashboard into "Team A".
  2. Create a second user with the Viewer role (Administration, then Users and access, then Users).
  3. Give that user View permission on "Team A" only, log in as them in a private browser window, and try to edit the dashboard.
the user can open the dashboard in "Team A" but cannot save changes. Then open the folder's Permissions tab and notice the default entries for the Viewer role: folder access starts from your organisation role, so tightening a folder means editing those defaults, not only adding a team.

Alerting: from a query to a notification

A dashboard only helps if someone is looking at it. Alerting makes Grafana look for you, and it is the part of the tool that has the most moving pieces, so take it slowly. The pieces fit together in a chain:

ALERT RULEquery + condition
→
ALERT INSTANCEone per series
→
NOTIFICATION POLICYroutes by labels
→
CONTACT POINTSlack, email, webhook

Grafana's alerting is built on the Prometheus Alertmanager, embedded in Grafana, and since version 11 it is the only alerting system; the older "legacy" alerting was removed.

Alert rules

An alert rule has four parts. First, one or more queries: the same kind of query you wrote for panels. Second, expressions that process the query results on the Grafana server; a common pair is a Reduce expression (collapse a time series to one number, such as the last value) and a Threshold expression (is that number above 0.5?). Third, a condition, which says which expression decides the outcome. Fourth, metadata: a folder, an evaluation group, labels, and annotations.

The evaluation group is a set of rules that are evaluated together at the same interval, for example every minute. The shortest interval the server allows is ten seconds by default, and there is rarely a reason to go near it; most rules do fine at one minute. The pending period says how long the condition must stay true before the alert actually fires. If you set it to 5 minutes, a spike that lasts 30 seconds does nothing. That single setting is the most effective way to stop alerts from flapping, and a great deal of alert fatigue comes from leaving it at zero.

An alert rule produces alert instances, one per series the query returns. If your query returns one series per job, then a single rule tracks each job separately. Each instance moves through states:

  • Normal: the condition is false.
  • Pending: the condition is true but the pending period has not elapsed yet.
  • Alerting (also called firing): the condition has held long enough, and a notification is sent.
  • Recovering: the condition has cleared but Grafana is holding the alert briefly before resolving it (added in Grafana 12).
  • No Data and Error: the query returned nothing, or failed.

The last two deserve attention. By default, when a rule gets no data, Grafana creates a synthetic alert named DatasourceNoData, and on a query failure it creates DatasourceError. Both carry the rule name and the data source's UID as labels. This is deliberate: a monitoring system that stays silent when it cannot read its data is worse than useless. In the rule's settings you can instead choose to treat no data as Normal, as Alerting, or to keep the last state. Choose carefully, since "Normal" means a broken exporter looks like a healthy system.

Labels and annotations

Labels identify and route: severity=critical, team=ml-platform, service=model-api. The notification policy reads labels to decide who is told. Annotations are human-readable text attached to the alert: a summary (one line), a description (what it means), and a runbook_url (what to do). A good annotation answers "what is wrong and what do I do first" for someone who was asleep a minute ago.

Annotations can use templates, which let them include the actual value:

TEXT
High error rate on {{ $labels.job }}: {{ printf "%.1f" $values.B.Value }} errors per second

The text inside {{ }} is Go template syntax; $labels holds the alert's labels and $values holds the results of the rule's expressions, referenced by the expression's name (here B).

Contact points

A contact point is a destination: email, Slack, Microsoft Teams, PagerDuty, Opsgenie, a generic webhook, Discord, Telegram, and more. Create one under Alerting, then Contact points. Every contact point has a Test button that sends a sample notification; always press it, because it catches wrong tokens and blocked networks immediately, instead of at three in the morning. A contact point can hold several integrations (email and Slack together). Email needs the [smtp] section of the configuration to be filled in, which is why email is often the first thing that does not arrive on a fresh install.

Notification policies

The notification policy tree decides which contact point receives which alerts. It starts with one default policy, which catches everything and sends it to a default contact point. You add nested policies with label matchers: "if severity = critical, send to the on-call pager; if team = ml-platform, send to their Slack channel". Alerts flow down the tree; the first matching child takes them, unless that child is marked to continue matching.

Policies also control timing through three settings with these defaults:

  • Group wait (30 seconds): how long to wait for more alerts of the same group before sending the first notification, so a burst arrives as one message.
  • Group interval (5 minutes): how long to wait before sending an update about a group that has new alerts.
  • Repeat interval (4 hours): how long to wait before reminding you about an alert that is still firing.

The group by setting chooses which labels split alerts into groups; grouping by alertname and job, say, sends one message per alert and job instead of one per instance. Those timings explain a common puzzle: a rule turns red on the dashboard but the Slack message arrives 30 seconds later. That is group wait working as designed, not a fault.

Silences and mute timings

Two tools stop notifications without touching the rule. A silence is a one-off, time-boxed mute matched by labels: use it during planned maintenance. A mute timing is a recurring schedule, such as weekends or a nightly batch window. Neither stops the rule from being evaluated; the alert still exists and still appears in the UI, but nobody is paged.

Building your first alert

Create an alert that fires when a scrape target is down. Open Alerting, then Alert rules, then New alert rule:

  1. Give the rule a name such as "Scrape target down".
  2. For the query, choose Prometheus and enter up{job="grafana"}.
  3. Add a Reduce expression (function Last) and a Threshold expression that fires when the reduced value is below 1. Set the threshold expression as the alert condition.
  4. Choose a folder and create an evaluation group with a one-minute interval.
  5. Set the pending period to 1 minute.
  6. Under labels, add severity=critical. Under annotations, add a summary such as "Grafana target is down".
  7. Choose the contact point, or leave routing to the notification policy.

To test it, stop the thing the rule watches. With Compose, stop Grafana's scrape target by editing prometheus.yml to point grafana:3000 at a port that is not open, and reload Prometheus, or simply pause a container that is scraped in your environment. Watch the rule's state change from Normal to Pending to Alerting in the UI, then confirm the notification arrives at your contact point.

Trap: an alert on the wrong shape of data. A rule can fail with a message like [sse.readDataError] [A] got error: input data must be a wide series but got type long. The expression stage expects time series data, and your query returned a table. The fix is to set the query's format to Time series, or to add a Reduce expression to convert the data into the shape the next step expects.
Shortcut. To experiment with contact points without Slack or email, create a webhook contact point pointing at a tiny local listener (or any request-inspection service you trust), then press Test. You can read the exact JSON Grafana sends, which shows you every label and annotation available to templates.

One more principle. Alert on symptoms that a person must act on, not on every number that moves. Grafana's own best-practice guidance suggests building alerts from the RED view of a service (Rate, Errors, Duration): how many requests, how many fail, how long they take. A high-CPU alert with no consequence trains people to ignore the pager, and an ignored pager is how real incidents get missed.

Try it
  1. Create a webhook contact point and press Test. Read the payload.
  2. Build the "Scrape target down" rule with a one-minute pending period.
  3. Break the scrape target on purpose and record the times: when the dashboard changes, when the rule goes Pending, when it goes Alerting, and when the notification arrives.
a gap of roughly the pending period plus group wait between "the thing broke" and "someone was told". Knowing those delays is what lets you tune them.

Configuration: files, environment variables, and precedence

Everything so far was done in the browser. The server itself, meanwhile, has settings (which port, where data lives, how users sign in, whether email works), and learning to change them is what turns Grafana from a toy into a service.

Grafana reads its settings from an ini file. Where that file is depends on the install: /etc/grafana/grafana.ini for Linux packages, /opt/homebrew/etc/grafana/grafana.ini for Homebrew on Apple Silicon, conf/custom.ini in a Windows or tarball install. Inside the install there is also a defaults.ini that lists every setting with its default; never edit that file, because upgrades overwrite it. Copy the lines you want to change into your own file. Comments start with ; or #, which matters because a line that still begins with ; is silently ignored, which is a classic source of "my change did nothing".

grafana.ini
[server]
http_port = 3000
root_url = https://grafana.example.com/

[security]
admin_user = admin
cookie_secure = true

[users]
allow_sign_up = false

[log]
mode = console file
level = info

[smtp]
enabled = true
host = smtp.example.com:587
user = alerts@example.com
password = $__env{SMTP_PASSWORD}
from_address = grafana@example.com

The file is divided into sections in square brackets, each holding key = value pairs. A few worth knowing on day one: [server] http_port (the port), [server] root_url (the public address, which must be right when you sit behind a proxy or sub-path, otherwise logins and links break), [users] allow_sign_up (off by default), [log] level (raise it to debug when troubleshooting), and [smtp] (required for email alerts).

Environment variables

Every setting can also be set with an environment variable named GF_<SECTION>_<KEY>, in upper case, with dots and dashes turned into underscores. So [server] http_port becomes GF_SERVER_HTTP_PORT, and [auth.google] client_secret becomes GF_AUTH_GOOGLE_CLIENT_SECRET. This is how you configure a container, where editing files is awkward:

BASH
docker run -d -p 3000:3000 --name=grafana \
  -e GF_SERVER_ROOT_URL=https://grafana.example.com/ \
  -e GF_USERS_ALLOW_SIGN_UP=false \
  -e GF_LOG_LEVEL=debug \
  --volume grafana-storage:/var/lib/grafana \
  grafana/grafana-enterprise:13.2.3

The precedence is fixed: built-in defaults, then the ini file, then environment variables, then command-line overrides. Later wins. If a setting "does not stick", an environment variable higher up the chain is usually overriding your file.

Keeping secrets out of the file

Never paste passwords into a file you commit. Grafana can read values from elsewhere. $__env{NAME} (used above for the SMTP password) inserts an environment variable. $__file{/path/to/file} inserts the contents of a file with the trailing newline trimmed, which suits Docker and Kubernetes secrets that are mounted as files. A HashiCorp Vault expander also exists, but it is an Enterprise feature. See the Vault guide for that side.

Some settings are worth setting deliberately from the start. [security] secret_key encrypts the secrets Grafana stores in its database (data source passwords, contact point tokens). If you ever lose it, those stored secrets become unreadable, so back it up together with the database. [paths] controls where data, logs, plugins, and provisioning files live. After editing the ini file, restart Grafana; it reads the file at start-up.

Trap: settings that only apply once. admin_password is read when the database is first created, then never again. Likewise a changed root_url needs a restart, and a mounted file that is not readable by the Grafana user is silently skipped. When a setting does not take effect, run with GF_LOG_LEVEL=debug, read the start-up log, and list the environment variables inside the container with docker exec grafana env to see which source won.
Try it
  1. Start the container with -e GF_LOG_LEVEL=debug and run docker logs grafana. Note how much more detail appears.
  2. Add -e GF_USERS_DEFAULT_THEME=light, recreate the container, and log in as a new user.
  3. Explain, in one sentence, why the theme setting changed the new user's page but not your existing profile.
the setting is only a default. Per-user preferences you saved earlier sit in the database and take priority over it.

Provisioning: configuration as files instead of clicks

Clicking is how you learn, but it is a poor way to run a service. If your only copy of a data source or dashboard is a row in a database that somebody set up by hand, you cannot review it, reproduce it on a second server, or rebuild it after a disaster. Provisioning fixes that: you describe data sources, dashboards, plugins, and alerting resources in files, and Grafana reads them at start-up.

This is the same idea as infrastructure as code (see the Terraform guide), applied to Grafana. The files live in a provisioning directory, which by default is /etc/grafana/provisioning on a package install, and is organised by kind: datasources/, dashboards/, plugins/, and alerting/.

Provisioning a data source

Here is the Prometheus data source from earlier, as a file:

provisioning/datasources/prometheus.yaml
apiVersion: 1

datasources:
  - name: Prometheus
    uid: prom-main
    type: prometheus
    access: proxy
    url: http://prometheus:9090
    isDefault: true
    editable: false
    jsonData:
      timeInterval: 15s

Every field has a reason. apiVersion: 1 is the file format version. uid: prom-main is a stable identifier that you choose. Always set it: if you leave it out, Grafana generates one, and then dashboards that refer to the data source by UID will not match when you recreate the server. access: proxy is the Server mode discussed earlier. isDefault marks the data source that new panels pick automatically. editable: false stops people from changing it in the UI, so the file stays the single source of truth. timeInterval tells Grafana the scrape interval so $__rate_interval can make good decisions. Secrets go under secureJsonData, and you should fill them from environment variables, for example httpHeaderValue1: $PROM_TOKEN, rather than typing them in the file. Environment interpolation works in values, not keys, and a literal dollar sign is written $$.

With Compose, mount the directory into the container:

docker-compose.yml
services:
  grafana:
    image: grafana/grafana-enterprise:13.2.3
    ports:
      - "3000:3000"
    volumes:
      - grafana_storage:/var/lib/grafana
      - ./provisioning:/etc/grafana/provisioning:ro

volumes:
  grafana_storage: {}

Recreate the container and the data source appears without you clicking anything. Delete the Grafana volume, start again, and it still appears; that is the point.

Provisioning dashboards

Dashboards use a two-part arrangement. A provider file tells Grafana where to look, and the dashboard JSON files sit in that folder.

provisioning/dashboards/default.yaml
apiVersion: 1

providers:
  - name: default
    orgId: 1
    folder: ""
    type: file
    disableDeletion: false
    allowUiUpdates: false
    updateIntervalSeconds: 30
    options:
      path: /var/lib/grafana/dashboards
      foldersFromFilesStructure: true

path is the directory with *.json dashboards. Grafana rescans it every updateIntervalSeconds, so changing a file changes the dashboard within about half a minute, with no restart. foldersFromFilesStructure: true turns sub-directories into folders (up to four levels deep). allowUiUpdates: false means people cannot save edits to these dashboards in the UI; they must change the file. That is the safe choice for a shared service, and it explains the error you will meet sooner or later.

Trap: "Cannot save provisioned dashboard". If you try to save a provisioned dashboard and Grafana refuses with a 400-level error saying the dashboard is provisioned, nothing is broken. The dashboard is owned by a file. Either edit the JSON file, or use Save as to make an editable copy, or (for experimentation only) set allowUiUpdates: true. A related error, "dashboard with the same uid already exists", means two files carry the same uid; give each dashboard a unique one.

How do you get the JSON? Build the dashboard in the UI, open Export (or Settings, then JSON Model), and copy it into a file. Make sure the dashboard has a stable uid and that its panels refer to your data source by the UID you provisioned (prom-main), not by a generated one. That pairing of "stable UIDs on both sides" is what makes a dashboard portable.

Other ways to manage it as code

Provisioning files are the simplest, but not the only, method. The Terraform provider for Grafana (published as grafana/grafana) manages folders, data sources, dashboards, and alert rules as Terraform resources, which suits teams already using Terraform. Git Sync, generally available since Grafana 13, connects a Git repository to a Grafana folder so dashboards stay in step in both directions: edit in the UI and Grafana proposes a commit, or merge a change in Git and Grafana picks it up. Kubernetes users often use the Helm chart's sidecar, which loads dashboards from ConfigMaps, or the Grafana Operator. You do not need any of these yet; just understand that "dashboards in version control" is the goal and there are several roads to it. The official dashboard best-practice guide makes the same point: keep dashboard JSON in version control and avoid drift.

Try it
  1. Create provisioning/datasources/prometheus.yaml with the file above and mount it as shown.
  2. Delete your Grafana volume (docker compose down -v) and start again.
  3. Check Connections, then Data sources.
  4. Export your dashboard JSON, save it into a dashboards directory mounted in the container, add the provider file, and restart.
the data source and the dashboard both reappear on a completely empty Grafana. You have turned a hand-built setup into something reproducible.

The HTTP API and service accounts

Anything you can do in the interface, a script can do through Grafana's HTTP API. That is how CI pipelines create dashboards, how scripts list what exists, and how other tools talk to Grafana. You do not need it on day one, but you will meet it quickly.

The health endpoint needs no authentication, as you already saw. Everything else needs a credential. There are two kinds.

Basic authentication uses a real user's name and password (-u admin:password). It is fine for a quick local experiment and required for a few server-admin tasks, but it ties a script to a person and is a poor habit in automation.

Service account tokens are the right tool. A service account is a non-human identity with an organisation role (Viewer, Editor, or Admin), and you attach one or more tokens to it. Tokens begin with glsa_, can be given an expiry, and can be revoked individually without touching anything else. They replace the older "API keys", which were migrated to service accounts automatically in earlier versions; if a tutorial says "create an API key", create a service account token instead.

Create one through the UI (Administration, then Users and access, then Service accounts), or from the command line with basic authentication:

BASH
curl -s -X POST -H "Content-Type: application/json" -u admin:admin \
  -d '{"name":"ci","role":"Editor"}' \
  http://localhost:3000/api/serviceaccounts

The response includes the new account's numeric id. Use it to mint a token with a thirty-day lifetime (2,592,000 seconds):

BASH
curl -s -X POST -H "Content-Type: application/json" -u admin:admin \
  -d '{"name":"ci-token","secondsToLive":2592000}' \
  http://localhost:3000/api/serviceaccounts/2/tokens

The response contains a key field with the token. It is shown once. Copy it into a password manager or a CI secret store immediately; Grafana cannot show it again. Replace the 2 with the id you received.

Now call the API with the token in an Authorization header:

BASH
export TOKEN="glsa_xxxxxxxxxxxx"

# search dashboards
curl -s -H "Authorization: Bearer $TOKEN" \
  "http://localhost:3000/api/search?query=overview"

# list data sources
curl -s -H "Authorization: Bearer $TOKEN" \
  http://localhost:3000/api/datasources

# fetch one dashboard by its UID
curl -s -H "Authorization: Bearer $TOKEN" \
  http://localhost:3000/api/dashboards/uid/<dashboard-uid>

Search returns a JSON list with each dashboard's uid, title, and folderTitle; the last call returns the full dashboard JSON, which is the same document you saw in JSON Model. A 401 means the token is missing, mistyped, or expired. A 403 means the token is valid but its role is too low for the action. A 404 usually means a wrong UID.

The API is in transition, and a beginner should know the shape of it. Grafana 12 introduced a Kubernetes-style dashboard API under /apis/dashboard.grafana.app/..., and Grafana 13 deprecated the older /api routes in favour of /apis. The old routes are not being switched off and remain fully working, but they will no longer receive updates. The new endpoints take a body with metadata (including the UID as name) and spec (the dashboard JSON). For learning, scripts, and quick automation, /api is perfectly fine today; for long-lived tooling that you are building now, check the current API reference in the Sources and prefer /apis where it covers your need. One related change is firm: the old numeric-ID data source endpoints are disabled by default, so always use UIDs.

Shortcut. Pipe output through jq to read it: curl -s -H "Authorization: Bearer $TOKEN" http://localhost:3000/api/search | jq '.[] | {uid, title}' shows just the two fields you care about.
Try it
  1. Create a service account with the Viewer role and a token with an expiry.
  2. Call /api/search with the token and find your dashboard's UID.
  3. Try to create a folder with the Viewer token, and note the status code. Then create the same token as Editor and try again.
a 403 for the Viewer token and success for the Editor token. The role on the service account, not your own login, decides what the script may do.

Plugins and the command line

Grafana ships with the common data sources and panels built in. A plugin adds anything else. There are three kinds: data source plugins (a new backend), panel plugins (a new visualisation), and app plugins (bundled pages and features). Plugins are signed: Grafana Labs signs its own, and community plugins can carry a community signature. By default Grafana refuses to load an unsigned plugin, which protects you from running unknown code inside the server. Note that backend plugins run as processes on the Grafana host, so install only what you trust.

Since Grafana 12, Angular plugins no longer load at all. If an old plugin logs a message about Angular not being supported, it must be rewritten by its author; there is no setting that brings it back.

The command line is the classic way to install plugins. Since Grafana 13 it is a subcommand of the single grafana binary:

BASH
grafana cli plugins list-remote
grafana cli plugins install grafana-clock-panel
grafana cli plugins ls
grafana cli plugins update grafana-clock-panel
grafana cli plugins remove grafana-clock-panel

list-remote lists what is available, install downloads a plugin, ls shows what is installed, and update and remove do what they say. Restart Grafana after changing plugins. The command works on the host that runs Grafana, because it reads and writes local files, and usually needs sudo. For a container, there is a neater way: set the environment variable GF_PLUGINS_PREINSTALL to a comma-separated list of id@version items, and Grafana installs them at start.

BASH
docker run -d -p 3000:3000 --name=grafana \
  -e GF_PLUGINS_PREINSTALL=grafana-clock-panel@1.0.1 \
  --volume grafana-storage:/var/lib/grafana \
  grafana/grafana-enterprise:13.2.3

The same command line also resets a forgotten admin password, as you saw earlier: grafana cli admin reset-admin-password 'NewPassword'. On a package install you may need to pass --homepath /usr/share/grafana so it finds its files.

Trap: the old command name. Scripts and tutorials that call grafana-cli or grafana-server as stand-alone commands fail on Grafana 13 because those commands were removed. Change them to grafana cli and grafana server. The systemd service is still called grafana-server, so systemctl restart grafana-server remains correct.
Try it
  1. Start a fresh container with GF_PLUGINS_PREINSTALL=grafana-clock-panel@1.0.1.
  2. Open Administration, then Plugins and data, then Plugins, and find the clock panel.
  3. Add a Clock panel to a dashboard.
the plugin is present without you running any install command by hand. This is the cleanest way to keep plugins reproducible.

Reading errors: the ones you will actually meet

Most early frustration with Grafana comes from a small set of failures. Learn to read them and most problems take a minute. Here are the common ones, grouped by where they appear. Always begin by asking two questions: did the request leave Grafana, and what did the other side say? The second answer is usually in the error text.

Data source errors

  • connection refused or dial tcp ... connect: connection refused when testing a data source. The address is wrong from Grafana's point of view. In Docker, use the service name and port, not localhost. Confirm by running docker exec grafana wget -qO- http://prometheus:9090/-/healthy or an equivalent from inside the Grafana container.
  • Bad Gateway in a panel. Grafana tried to reach the backend and could not. Same causes as above.
  • client_error: client error: 401. The backend wants credentials. Add them in the data source's authentication section.
  • 422 ... parse error. Your PromQL (or SQL) has a syntax mistake. The message points at the position.
  • context deadline exceeded or Timeout exceeded. The query is too heavy. Shorten the time range, reduce the series, or raise the data source's query timeout.
  • No data with no error. The query is valid but returned nothing: check the time range, the label values, and any variable selection.

Login and access errors

  • Login failed or Invalid username or password. Wrong credentials, or the admin password from the environment was ignored because the database already existed. Use grafana cli admin reset-admin-password.
  • Origin not allowed on login behind a proxy. Grafana compares the request's origin to root_url; set root_url to the public address and make the proxy forward the Host header.
  • 401 Unauthorized or Invalid API key from the API. Use a service account token that starts with glsa_.
  • 403 Forbidden from the API or the UI. The role is too low for that action.

Start-up and storage errors

  • mkdir /var/lib/grafana: permission denied or unable to open database file. Grafana cannot write to its data directory. Fix ownership with chown -R grafana:grafana /var/lib/grafana, or in Docker run with a user that owns the bind mount (--user "$(id -u)"). The container's default user is 472, which matters for Kubernetes fsGroup.
  • database is locked. You are using SQLite with concurrent writers. The fix is to move to PostgreSQL or MySQL, and never run two Grafana processes on one SQLite file.
  • Dashboards vanish after restart. The data directory was not on a volume, so the container's database was thrown away with the container. Attach a volume at /var/lib/grafana.

Plugin and provisioning errors

  • Plugin ... is not signed or an invalid-signature message. Grafana refuses unsigned plugins by default. Install a signed version; for a plugin you wrote yourself, allow it with [plugins] allow_loading_unsigned_plugins = plugin-id and understand that you are accepting the risk.
  • Panel plugin not found. The plugin is not installed or did not load; install it and restart.
  • YAML errors in provisioning such as yaml: line 7: did not find expected key. Indentation. YAML allows spaces only, never tabs, and a misaligned line shifts everything under it.
  • the same UID is used more than once. Make every provisioned dashboard and data source UID unique.
Shortcut. The fastest diagnosis tool is the log. Run docker logs grafana (or journalctl -u grafana-server) while you reproduce the problem, and read the last few lines. Raising the level with GF_LOG_LEVEL=debug shows the request Grafana sent and the reply it got. For a failing panel, the Query inspector (panel menu, then Inspect, then Query) shows the same thing without touching the server.
Try it
  1. Break a data source on purpose: change its URL to a port nothing listens on, and press Save and test.
  2. Open a panel that uses it, then open Inspect and read the error.
  3. Find the same failure in docker logs grafana.
the same root cause appears three times in three places: the test banner, the Inspect panel, and the log. Learn to triangulate between them.

Putting it all together

Now combine everything into one small project that you could hand to a colleague: a monitoring stack with Grafana and Prometheus that starts from nothing with one command, arrives with its data source and a dashboard already in place, and has an alert ready to use. You have already built every piece; this exercise is about seeing them as a whole.

The folder layout is the first design decision. Keep everything in one Git repository:

TEXT
monitoring/
  docker-compose.yml
  prometheus.yml
  provisioning/
    datasources/
      prometheus.yaml
    dashboards/
      default.yaml
  dashboards/
    overview.json

Follow these steps:

  1. Write the Compose file with Grafana (pinned to 13.2.3) and Prometheus (pinned to an exact tag you choose). Give both named volumes. Mount ./prometheus.yml into Prometheus, and mount ./provisioning and ./dashboards into Grafana (the dashboards at /var/lib/grafana/dashboards, matching the provider file).
  2. Write prometheus.yml so Prometheus scrapes itself and Grafana, and set GF_METRICS_ENABLED=true on Grafana.
  3. Provision the data source with a stable UID (prom-main), access: proxy, and editable: false.
  4. Provision the dashboard provider, pointing at the mounted directory.
  5. Export your dashboard from the UI, making sure panels use the prom-main UID and the dashboard has its own stable uid. Save it as dashboards/overview.json. It should contain the request-rate time series, the healthy-targets Stat, the table, and the job variable.
  6. Take the admin password from the environment with GF_SECURITY_ADMIN_PASSWORD read from a .env file that you do not commit, so no secret sits in the repository.
  7. Bring it up with docker compose up -d, then verify: curl -s http://localhost:3000/api/health shows "database": "ok"; the data source test is green; the dashboard opens with data.
  8. Add the alert in the UI ("Scrape target down", one-minute pending period, severity=critical) and a webhook contact point, and test both.
  9. Destroy and recreate with docker compose down -v followed by docker compose up -d. Everything you provisioned returns; only the alert rule and contact point, which you created by hand, do not.

That last observation is the key lesson, and a good way to finish: whatever you did by clicking is fragile, and whatever you put in a file survives. The natural next step is to provision the alerting resources too (the provisioning/alerting directory takes contact points, policies, and rule groups), or manage them with Terraform. The same instinct applies to the dashboards: the JSON in dashboards/ should be edited, reviewed, and committed like code, instead of being changed in the browser and forgotten.

For a service you would share, add a short checklist before calling it done:

  1. Versions are pinned: no latest anywhere.
  2. The admin password is changed and is not in the repository.
  3. Data source credentials are read-only and scoped.
  4. Grafana's own data is on a volume; you know where the database is.
  5. Dashboards and data sources come from files with stable UIDs.
  6. At least one alert has been tested end to end, including the notification.
  7. Anonymous access is off and sign-up is off.
Try it
  1. Build the whole repository above from scratch in an empty folder.
  2. Push it to a Git repository and clone it into a second folder.
  3. Run docker compose up -d there and confirm the dashboard appears with data.
a colleague-ready monitoring stack that goes from clone to working dashboard in one command, which is the standard a real team holds itself to.

What you can now do, and what comes next

You started without knowing what Grafana was, and you can now do real work with it. You can explain that Grafana queries other systems and stores only its own bookkeeping. You can name the five nouns (data source, query, panel, dashboard, alert rule) and say which layer to look at when a chart is wrong. You can install Grafana in a container or a package, verify it with the health endpoint, and read its logs. You can connect Prometheus and a SQL database, run PromQL and SQL in Explore, and build a dashboard with time series, stat, and table panels. You can add variables that make a dashboard reusable. You can build an alert that goes from rule through notification policy to a contact point, and you know why the message arrives after a short delay. You can configure the server with an ini file and environment variables, keep secrets out of files, put data sources and dashboards in version control with provisioning, call the API with a service account token, and install plugins without the removed commands. And you know the errors you will meet first and where each one shows up.

Some habits to carry forward. Pin versions and upgrade deliberately, reading the "what's new" page each time, since Grafana changes quickly. Put everything you can into files. Give Grafana read-only credentials, because every user of an organisation can query every data source in it. Give every panel a title and a unit. Alert on symptoms someone must act on, with a pending period, and always test the notification.

What the next levels cover:

  • Mid-level goes under the hood of queries and dashboards: moving off SQLite to PostgreSQL, single sign-on, folder and team design, recording rules and alert quality, Grafana with Kubernetes and Helm, templated dashboards at scale, and Git Sync and Terraform workflows.
  • Senior treats Grafana as a platform: high availability, the alerting cluster, security and secrets, multi-tenancy, upgrades and migrations (including the 12 to 13 checklist), cost control, and knowing when a different tool fits better.

Natural neighbours in this catalogue: the Prometheus guide for the metrics store behind most dashboards, the OpenTelemetry guide for producing the telemetry, the Docker guide and Kubernetes guide for running it all, Helm for installing Grafana on a cluster, and the Evidently guide for model-quality metrics that you can chart in Grafana.

Sources