تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
ELK StackDevOpsObservability3 مستويات108 قسمًايغطّي Elastic Stack 9.5دليل بالإنجليزية

The Complete ELK Stack Guide

Search and analyse logs with Elasticsearch, Logstash and Kibana. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
17sections
32examples

This is part one of three. It covers everything you need to do real work with the ELK stack, not a teaser. By the end you can start a working stack on your laptop, put documents into Elasticsearch, search and summarise them, build a dashboard in Kibana, ship a real log file into the stack with Filebeat, and reshape messy log lines with Logstash. You will also be able to read the ten errors that account for most of a beginner's pain. The Mid-level and Senior guides take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and this material only sticks once you have watched your own data arrive, get searched, and occasionally get rejected.

The guide is written against Elastic Stack 9.5.4, released on 15 September 2026. That number matters for a practical reason: 9.5.4 and 9.4.7 fixed a denial-of-service vulnerability in Elasticsearch (ESA-2026-182, CVE-2026-94396) that affects 9.2.0 through 9.4.6 and 9.5.0 through 9.5.3, and Elastic lists no workaround other than upgrading. Whenever you install the stack, install 9.5.4 or later.

What the ELK stack is, and the problem it solves

Every running system writes down what it is doing. A web server records each request, an application records each error, a database records each slow query, a load balancer records each connection. Those records are logs, and on one machine you can read them with tail -f and grep. That stops working the moment you have ten machines, or one machine producing a million lines an hour, or a bug report that says "something was slow at about three in the morning last Tuesday".

The ELK stack exists to answer three questions about that mountain of text. Where is all of it? It gets collected from every machine into one place. How do I find the one line I care about? It gets indexed, so a search across a billion lines returns in milliseconds. What is the shape of it? It can be counted, grouped and charted, so you can see that errors tripled at 03:10 and that all of them came from one host.

SOURCESservers, apps, containers
→
SHIPPERFilebeat, Elastic Agent
→
LOGSTASHoptional: parse and route
→
ELASTICSEARCHstore and search
→
KIBANAexplore and chart

That diagram is the whole stack. Read it left to right and you have followed one log line from the moment an application wrote it to the moment a person saw it on a screen.

Before tools like this, the usual answer was a pile of scripts. Someone would rsync log files to a central server nightly, run grep over them, and paste the results into a ticket. Some teams loaded logs into a relational database, which works until the volume grows and every query becomes a full-table scan. Others bought a proprietary log-management product priced by the gigabyte. The stack that became ELK was the open, searchable alternative, and it turned out to be useful well beyond logs: the same engine now powers website search, product catalogues, security analytics, and metrics dashboards.

Two ideas explain why it is fast, and it is worth meeting them now because they shape every command later.

Search is built on an inverted index. A book has an index at the back: look up a word, get the page numbers. Elasticsearch builds the same thing for every word in every document, so finding all documents containing timeout is a lookup rather than a scan. That is why searching a billion lines feels instant, and also why how a field is indexed decides what you can do with it.

Data is spread over shards. An index is split into pieces called shards that can live on different machines, so the work of searching is shared and the data can grow beyond one disk. On your laptop you will have one machine and the machinery will mostly be invisible, but the words "shard" and "replica" will appear in health output, so we define them properly soon.

What people use it for:

📜

Central logging

Every service writes logs; they all land in one searchable place, with retention you control.

🔎

Application search

The search box on a website or app: full-text matching, typo tolerance, filters and facets.

📈

Metrics and dashboards

CPU, latency and request counts over time, charted in Kibana alongside the logs that explain them.

🛡️

Security analytics

Authentication events, network flows and audit trails searched and correlated to find an intruder.

For MLOps work specifically, this is where model-serving logs go. When a prediction endpoint starts returning errors, or a feature pipeline stalls, the evidence is in logs and metrics from five different services. A stack that lets you search them together is often the difference between a ten-minute diagnosis and a two-day one. A note for readers working for Gulf or Egyptian employers: logs routinely contain user identifiers and IP addresses, so where the cluster physically lives can be a data-residency question. Self-managed clusters and the regional options of the cloud providers are both possible; decide that before you ship production logs anywhere.

Try it
  1. Pick an application you have run, and find where it writes its logs.
  2. Answer this question using only grep and tail: how many errors happened in the last hour?
  3. Now imagine the same application on twenty servers, and answer the same question.
the first answer takes a minute; the second is the moment you realise why centralised, indexed logs exist.

The names: ELK, the Elastic Stack, Beats and Elastic Agent

The vocabulary around this stack confuses almost everybody, because the name changed while the products multiplied. Sort it out once and the documentation stops being baffling.

ELK is an acronym for three products: Elasticsearch, Logstash and Kibana. It is what most people, and most job advertisements, still say. Elastic Stack is the official name, and it includes the three ELK products plus the collection side: Beats and Elastic Agent. You will see both terms, and they describe the same ecosystem. In this guide "ELK" means the three core products and "the stack" means everything.

Here is each piece in one sentence, with the port it listens on, because ports are how you will check that things are running:

  • Elasticsearch is the distributed search and analytics engine, built on the Apache Lucene library. It stores your data and answers queries through a JSON REST API on port 9200. Port 9300 is used between Elasticsearch nodes and is not something you call.
  • Kibana is the web interface on port 5601. It is where you search, build dashboards, manage the cluster and run commands in a console.
  • Logstash is a data-processing pipeline that reads events from inputs, transforms them with filters, and writes them to outputs such as Elasticsearch. Its monitoring API is on port 9600.
  • Beats are small, single-purpose shippers written in Go. Filebeat ships log files; there are others for metrics, network data, uptime checks and Windows event logs.
  • Elastic Agent is a single unified agent that does the work of several Beats and is managed centrally from Kibana through Fleet. Elastic recommends it for new deployments. Beats are still supported and still shipped, and they remain the simplest thing to learn first, which is why this guide uses Filebeat in its hands-on section.

Two more terms appear constantly. An integration is a packaged recipe for one data source, for example nginx or AWS: the collection settings, the parsing, the field mappings and ready-made dashboards bundled together. ECS, the Elastic Common Schema, is an agreed set of field names (@timestamp, host.name, log.level, source.ip, message) so that logs from different sources can be searched with the same names.

Do not learn retired features Older tutorials and answers online describe things that are gone. Elastic's Enterprise Search products (App Search and Workplace Search) were removed in version 9.0. The Logstash "modules" framework was removed in 9.0. Mapping types (the old /index/type/id URLs) were removed in 8.0. If a tutorial uses any of these, it is out of date; close the tab.

There is also a licensing point that employers sometimes ask about. The default download is under the Elastic License 2.0. Since September 2024 the source is additionally available under the AGPLv3, an OSI-approved open-source licence, so the stack is open source again if you choose that option. The free "Basic" tier already includes security (TLS, users and roles, API keys), Kibana, Fleet and the ES|QL query language, which is everything this guide uses. Some features, such as machine-learning anomaly detection, single sign-on for self-managed clusters and searchable snapshots, need a paid subscription. The local installer used below gives you a 30-day trial of those and then falls back to Basic.

Try it
  1. Write the five boxes from the flow diagram on paper.
  2. Under each, write the product that fills it, its port if it has one, and one word for its job.
  3. Circle the box that is optional.
Logstash is the circled one. Many pipelines go shipper straight to Elasticsearch, and you add Logstash when you need heavier parsing, enrichment or routing.

The mental model: the nouns you need

You can run this stack for a year without understanding Lucene internals. You cannot run it for a day without a handful of nouns. Learn these, and the error messages will read like sentences.

A document is one JSON object, the unit of storage: one log line, one product, one user. Every document has an _id and a _source, which is the original JSON you sent.

An index is a named collection of documents that share a purpose, like web-logs or products. If you know relational databases, an index is roughly a table, and a document is roughly a row, with the important difference that the documents are JSON and do not all have to look alike.

A field is one key inside a document. The mapping is the schema: it says what type each field has. The types you will meet first are text, keyword, long, double, date, boolean, and ip. The text and keyword pair is the most important distinction in the whole stack:

  • A text field is analysed: broken into lowercase words, so a search for timeout finds "Connection Timeout after 30s". It is for full-text search and cannot be sorted or grouped efficiently.
  • A keyword field is stored as one exact value. It is for filters, sorting and counting: status codes, host names, log levels, user ids.

Dynamic mapping means that if you send a field Elasticsearch has not seen, it guesses a type. A JSON string it guesses as text and adds a .keyword sub-field, so host becomes both a searchable host and an exact host.keyword. That convenience is wonderful on day one and a source of trouble on day thirty, which is why we look at explicit mappings later.

A shard is one piece of an index. Internally, each shard is a complete Lucene index. There are two kinds: primary shards hold the original data, and replica shards are copies placed on different nodes for resilience and read speed. The number of primary shards is fixed when you create the index; the number of replicas can change any time.

A node is one running Elasticsearch process, and a cluster is one or more nodes sharing a cluster.name. On your laptop you have a cluster of one node.

Cluster health is reported as a colour, and you will check it constantly:

Colour Meaning
green Every primary and every replica shard is assigned to a node
yellow Every primary is assigned, but some replicas are not
red At least one primary shard is not assigned, so some data is unavailable

The surprise that catches every beginner: a single-node cluster with an index that asks for one replica is yellow, forever. A replica may not live on the same node as its primary (a copy on the same machine would protect against nothing), and there is no second node, so the replica has nowhere to go. Yellow on a laptop is normal, not a fault.

Two more nouns round out the set. A data view is a Kibana object that says "these indices (or data streams) are the ones I want to explore"; older material calls it an index pattern. A data stream is an append-only series of hidden indices behind one name, designed for logs and metrics; the shippers in this guide write to data streams, and you will see names like logs-nginx.access-default. Its naming convention is <type>-<dataset>-<namespace>.

Finally, the near-real-time rule. A document you index is not searchable instantly; it becomes visible after the next refresh, which by default happens every second. So if you index and immediately search, you may miss your own document for a moment. This is by design, and it is the reason Elasticsearch is called near-real-time rather than real-time.

text field

  • Analysed into lowercase words
  • Searched with match
  • Good for messages and descriptions
  • Cannot sort or group on it

keyword field

  • Stored as one exact value
  • Searched with term
  • Good for status, host, level, id
  • Sorts and groups fine
Try it
  1. For each field, decide text or keyword: message ("user login failed"), http.status (500), host.name ("web-01"), product.description.
  2. Explain in one sentence why a one-node cluster can never be green with one replica.
message and description are text; status and host are keyword. The one-node answer is "a replica cannot share a node with its primary".

Installing the stack on your laptop

There are several ways to install, and choosing well saves an afternoon. For learning, use start-local: a script from Elastic that uses Docker Compose to start Elasticsearch and Kibana with sensible settings in one command. It works on macOS, Linux, and Windows through WSL. For a real server you use the operating-system packages or a container image, and we cover those briefly below so you recognise them. If you are new to containers, the Docker guide explains the ideas this section leans on, and it is worth reading first.

You need Docker running, and Docker Compose version 2 (it ships with current Docker Desktop). Then:

BASH
curl -fsSL https://elastic.co/start-local | sh

The script downloads a small project and starts it. When it finishes you have a new folder called elastic-start-local/ containing a Compose file, a hidden .env file, and scripts named start.sh, stop.sh and uninstall.sh. Elasticsearch is at http://localhost:9200 and Kibana at http://localhost:5601.

The .env file is the part to understand. It holds two secrets the script generated for you: ES_LOCAL_PASSWORD, the password for the built-in superuser called elastic, and ES_LOCAL_API_KEY, an API key. Look at them with:

BASH
cd elastic-start-local
grep ES_LOCAL_ .env

Two flags are useful. -v <version> pins a version, and --esonly skips Kibana. Because this guide is written for 9.5.4, you can pin it explicitly, but the default already installs a current release:

BASH
curl -fsSL https://elastic.co/start-local | sh -s -- -v 9.5.4
start-local is for your laptop only It runs over plain HTTP with password authentication, with no TLS. That is deliberate, because it keeps the learning curve flat, but it means you must never expose those ports to a network or use this setup for real data. A production cluster uses HTTPS on both the client and the node-to-node channel, as described next.

Installing on a real machine

On a server you install packages or run an image. You do not need to memorise these steps, but you should know what they look like and what they print. On Debian or Ubuntu, add Elastic's signing key and repository, then install each product from the same repository:

BASH
wget -qO - https://artifacts.elastic.co/GPG-KEY-elasticsearch | sudo gpg --dearmor -o /usr/share/keyrings/elasticsearch-keyring.gpg
sudo apt-get install apt-transport-https
echo "deb [signed-by=/usr/share/keyrings/elasticsearch-keyring.gpg] https://artifacts.elastic.co/packages/9.x/apt stable main" | sudo tee /etc/apt/sources.list.d/elastic-9.x.list
sudo apt-get update && sudo apt-get install elasticsearch
sudo /bin/systemctl daemon-reload
sudo /bin/systemctl enable elasticsearch.service
sudo systemctl start elasticsearch.service

Read the terminal output of that install. The package prints the automatically generated password for the elastic user exactly once. If you scroll past it, you will need to generate a new one later with elasticsearch-reset-password. Kibana and Logstash install the same way: sudo apt-get install kibana and sudo apt-get install logstash. On Red Hat-family systems the equivalent uses an .repo file and dnf. On macOS or Linux you can also download a .tar.gz archive, extract it, and run ./bin/elasticsearch, which runs in the foreground and prints the password and a Kibana enrollment token to your terminal. On macOS the bundled Java runtime may be blocked by Gatekeeper, and the documented fix is xattr -d -r com.apple.quarantine . inside the extracted folder. Homebrew is not recommended for Elasticsearch: Elastic no longer maintains a tap, and the community formula lags the current release.

On Windows you download the .zip, unzip it, and run .\bin\elasticsearch.bat, or install it as a service with .\bin\elasticsearch-service.bat install.

With Docker alone, without Compose, the official quick start creates a network, runs the image, and resets the password:

BASH
docker network create elastic
docker run --name es01 --net elastic -p 9200:9200 -it -m 1GB docker.elastic.co/elasticsearch/elasticsearch:9.5.4
docker exec -it es01 /usr/share/elasticsearch/bin/elasticsearch-reset-password -u elastic
docker exec -it es01 /usr/share/elasticsearch/bin/elasticsearch-create-enrollment-token -s kibana
docker run --name kib01 --net elastic -p 5601:5601 docker.elastic.co/kibana/kibana:9.5.4

Two facts about Elasticsearch that surprise people. First, it bundles its own Java, so you do not install Java for it and should not point it at another. Second, every component in one deployment should run the same version: Elasticsearch 9.5.4 with Kibana 9.5.4 and Logstash 9.5.4.

What first-start security does

Since version 8, security is on by default. When a brand-new node starts for the first time, it configures itself: it generates a certificate authority and TLS certificates for both the HTTP and the node-to-node channels, turns on authentication, sets a password for the elastic user, and prints a Kibana enrollment token that is valid for 30 minutes. The certificate for the HTTP channel lives in certs/http_ca.crt under the config directory, and this file is what clients need to trust the server. Package installs put it at /etc/elasticsearch/certs/http_ca.crt.

Connecting Kibana to such a node takes four steps: start Kibana, open the URL it prints (something like http://localhost:5601/?code=123456), paste the enrollment token, then enter the six-digit verification code that bin/kibana-verification-code prints. Then log in as elastic. If the token expired, make a new one with bin/elasticsearch-create-enrollment-token -s kibana.

Auto-configuration is skipped in some situations, all of which are worth knowing because they explain "why did it not generate a password": the data directory was not empty, security settings were already in elasticsearch.yml, standard output was redirected to a file, or discovery settings for a multi-node cluster were already configured.

Try it
  1. Run the start-local command above.
  2. Open http://localhost:5601 in a browser and log in as elastic with the password from .env.
  3. Run ./stop.sh inside the folder, then ./start.sh, and confirm that Kibana comes back.
a Kibana home page after login. Start-up takes a minute or two on a cold machine; "Kibana server is not ready yet" during that time is normal.

Checking that the installation works

Never assume a service is healthy because the install command ended without an error. Elasticsearch has a REST interface, so the test is a plain HTTP request, and the habit of checking is worth building now.

With start-local, the password is in .env. Load it into a shell variable so you never paste it into history:

BASH
export ELASTIC_PASSWORD=$(grep ES_LOCAL_PASSWORD .env | cut -d= -f2)
curl -u elastic:$ELASTIC_PASSWORD http://localhost:9200

The response is a JSON document describing the node. The parts to read:

JSON
{
  "name": "es01",
  "cluster_name": "docker-cluster",
  "cluster_uuid": "k3Xp0s7fT9qH2mVwYl8aBg",
  "version": { "number": "9.5.4" },
  "tagline": "You Know, for Search"
}

Your names and uuid will differ. What matters is version.number, which must match what you meant to install, and the tagline, which proves the request reached Elasticsearch rather than something else on that port.

On a package or archive install, security is on with TLS, so the same check uses https and the CA certificate file:

BASH
curl --cacert /etc/elasticsearch/certs/http_ca.crt -u elastic:$ELASTIC_PASSWORD https://localhost:9200

If you forget --cacert, curl stops with SSL certificate problem: self-signed certificate in certificate chain. That is curl correctly refusing to trust a certificate authority it has never heard of. The fix is to provide the certificate, not to add -k, which switches verification off and teaches a bad habit.

Now ask the cluster how it feels:

BASH
curl -u elastic:$ELASTIC_PASSWORD "http://localhost:9200/_cluster/health?pretty"

Read the status field: green, yellow or red, as defined earlier. Also look at number_of_nodes (one on your laptop) and unassigned_shards. On a fresh single-node cluster with no user indices, you should see green. After you create your first index with a replica, expect yellow.

For the other components, the checks are equally direct. Kibana's health is at http://localhost:5601/api/status, or simply the login page loading. Logstash has a smoke test that needs no Elasticsearch at all:

BASH
bin/logstash -e 'input { stdin { } } output { stdout {} }'

Type a line, press Enter, and Logstash prints it back as an event. Its monitoring API, once running, answers curl 'localhost:9600/?pretty'. For Filebeat the two checks are filebeat test config -e, which validates the YAML, and filebeat test output, which tests the connection to Elasticsearch. A third, for Elastic Agent, is sudo elastic-agent status.

Use the console, not curl, for exploring curl is right for scripts and for first checks. For everything exploratory, use Kibana Dev Tools (the menu, then Dev Tools, then Console). You type GET _cluster/health with no host, no quotes and no password, and it runs as your logged-in user. The rest of this guide shows requests in console syntax for that reason.
Try it
  1. Run the curl command against port 9200 and find version.number in the output.
  2. Run the cluster health request and note the status and number_of_nodes.
  3. Open Dev Tools in Kibana and run GET / and GET _cluster/health.
the same data three ways. Seeing the console return what curl returned is the moment Dev Tools stops being mysterious.

Your first index: putting data in and getting it out

Everything so far has been setup. This section is the first time you actually store something. Open Dev Tools; all the requests below are in its syntax. The left pane is where you type, the right pane is the response, and the green triangle (or Ctrl+Enter) runs the request under your cursor.

Start by indexing one document. You could create the index first, but Elasticsearch will create it for you on first write, with a guessed mapping:

HTTP
PUT web-logs/_doc/1
{
  "@timestamp": "2026-09-30T10:00:00Z",
  "host": "web-01",
  "level": "error",
  "status": 500,
  "message": "Connection timeout while calling payment service"
}

The URL reads as: in the index web-logs, create document id 1. The word _doc is fixed; it is the name of the endpoint, not a type you choose. The response includes "result": "created" and a _version of 1. Run the same request again and you get "result": "updated" and version 2, because the same id replaced the document. To let Elasticsearch generate the id, use POST web-logs/_doc with no id.

Read a document back by id, and then change part of it:

HTTP
GET web-logs/_doc/1

POST web-logs/_update/1
{ "doc": { "level": "critical" } }

DELETE web-logs/_doc/1

Real work involves many documents at once, and one request per document is painfully slow. The bulk API sends many operations in a single request. Its body is unusual: newline-delimited JSON, alternating an action line and a document line, and it must end with a newline:

HTTP
POST _bulk
{ "index": { "_index": "web-logs" } }
{ "@timestamp": "2026-09-30T10:00:05Z", "host": "web-01", "level": "info", "status": 200, "message": "GET /home served in 34 ms" }
{ "index": { "_index": "web-logs" } }
{ "@timestamp": "2026-09-30T10:00:09Z", "host": "web-02", "level": "error", "status": 502, "message": "Bad gateway from upstream model-server" }
{ "index": { "_index": "web-logs" } }
{ "@timestamp": "2026-09-30T10:01:14Z", "host": "web-01", "level": "warn", "status": 429, "message": "Rate limit reached for client 203.0.113.7" }
{ "index": { "_index": "web-logs" } }
{ "@timestamp": "2026-09-30T10:02:30Z", "host": "web-03", "level": "error", "status": 500, "message": "Connection timeout while calling payment service" }
{ "index": { "_index": "web-logs" } }
{ "@timestamp": "2026-09-30T10:03:41Z", "host": "web-02", "level": "info", "status": 200, "message": "GET /pricing served in 51 ms" }

The response lists one result per line and has a top-level "errors": false if everything worked. Always check that flag. A bulk request that partly fails still returns HTTP 200, with "errors": true and the failing items marked inside, so a script that only checks the status code will silently lose data.

Now look at what you created. Two requests from the _cat family give human-readable tables, and ?v adds column headers:

HTTP
GET _cat/indices?v
GET web-logs/_count
GET web-logs/_mapping

The first shows your index with its health (yellow on one node, as explained), primary and replica counts (1 and 1 by default), and document count. The mapping shows what Elasticsearch guessed. You will see @timestamp as date, status as long, and host, level and message as text each with a .keyword sub-field. That is dynamic mapping in action.

Try it
  1. Index the single document with id 1, then the five-document bulk.
  2. Run GET _cat/indices?v and confirm web-logs reports six documents (id 1 plus five from the bulk).
  3. Run GET web-logs/_mapping and find the .keyword sub-field on host.
docs.count of 6, a yellow health, and a mapping you did not write. If you immediately saw 0 docs, wait a second; that is the one-second refresh interval.

Searching: the query language you will use every day

Searching uses GET index/_search with a JSON body called the Query DSL. With no body, it returns the first ten documents. The shape of every response is the same: a hits object containing total (how many matched) and hits (the first ten, each with _score and _source).

Full-text search with match

HTTP
GET web-logs/_search
{
  "query": { "match": { "message": "timeout" } }
}

match is for text fields. It analyses your search text the same way the field was analysed at index time, which is why lowercase timeout finds "Connection timeout while calling payment service". Each hit has a _score, a relevance number: documents where the term is rarer or more prominent score higher. Try "message": "payment timeout", and you get documents with either word, ranked with both-word matches first.

Exact search with term, and filtering with range

For exact values, use term against a keyword field. Because dynamic mapping made host.keyword, that is the one to search:

HTTP
GET web-logs/_search
{
  "query": { "term": { "host.keyword": "web-01" } }
}

A classic mistake is term on the analysed host field with capital letters; the stored token is lowercase, so nothing matches. The rule of thumb: match for text, term for keyword, numbers and dates.

range works on numbers and dates. Here is every response with a status of 500 or above:

HTTP
GET web-logs/_search
{
  "query": { "range": { "status": { "gte": 500 } } }
}

Combining conditions with bool

Real questions have several parts: "errors from web-01 in the last hour whose message mentions timeout". The bool query combines clauses, and the four names are worth memorising:

Clause Meaning Affects score?
must Every clause must match Yes
filter Every clause must match, yes or no No, and results are cached
should Matching is a bonus (or required, if nothing else is) Yes
must_not No clause may match No
HTTP
GET web-logs/_search
{
  "query": {
    "bool": {
      "must":   [ { "match": { "message": "timeout" } } ],
      "filter": [
        { "term":  { "level.keyword": "error" } },
        { "range": { "@timestamp": { "gte": "now-1d/d" } } }
      ]
    }
  },
  "sort": [ { "@timestamp": "desc" } ],
  "size": 20,
  "_source": ["@timestamp", "host", "message"]
}

Notice the division of labour: the relevance question ("is about timeouts") sits in must, and the yes/no conditions sit in filter, which is faster because Elasticsearch does not compute a score for them and can cache the result. sort orders by time, size returns twenty rows instead of ten, and _source limits which fields come back. A good habit: any condition that is not about relevance belongs in filter.

Two limits to know now. Results are capped: from plus size cannot exceed 10,000, and deeper paging needs a different technique covered in the Mid-level guide. And a search response has "timed_out" and "_shards" sections; if _shards.failed is above zero, the results are partial.

A search that returns nothing is not always an empty index Before you conclude the data is missing, check three things in order: the index name or data view matches, the time range (Kibana defaults to the last 15 minutes, and your sample data is from September 2026), and whether you used term on a text field. Those three cause nearly every "my search returns nothing" report.
Try it
  1. Find all documents whose message mentions gateway.
  2. Find all documents from web-02 using term.
  3. Write a bool query for level.keyword equal to error with status above 499, sorted newest first.
one hit for gateway, two for web-02, and three errors for the bool (the two from the bulk plus document 1, if you kept it). If term returned zero, check that you used the .keyword field.

Mappings: taking control of the schema

Dynamic mapping got you started, but it makes decisions for you, and some of them are wrong for logs. Every string becomes both a text field and a .keyword sub-field, which doubles the storage for fields like host that you only ever filter on. A number that arrives as a string in one document ("status": "200") fixes the type for the whole index; a later document with "status": "unknown" is then rejected. An unexpected field name that appears in one application's logs silently becomes a permanent column. None of these are bugs; they are the price of "it just works".

The alternative is to declare the mapping before you write data. You create the index with the schema you want:

HTTP
PUT app-logs
{
  "mappings": {
    "properties": {
      "@timestamp": { "type": "date" },
      "host":       { "type": "keyword" },
      "level":      { "type": "keyword" },
      "status":     { "type": "integer" },
      "client_ip":  { "type": "ip" },
      "message":    { "type": "text" }
    }
  }
}

Each field now has exactly the type you chose. host and level are pure keywords, so term queries and aggregations work on the field name itself, with no .keyword suffix. client_ip is an ip type, which lets you search by CIDR range such as 203.0.113.0/24. message is full-text only, which is all a log message needs.

You can also set how strict the index is about fields you did not declare, through the dynamic setting in the mapping:

Value Behaviour
true New fields are added automatically (the default)
false New fields are stored in _source but not indexed, so not searchable
strict A document with an undeclared field is rejected
runtime New fields become runtime fields, computed at query time

The rule that matters most: you cannot change the type of an existing field. If status was mapped as text and you want integer, the fix is to create a new index with the right mapping and copy the data across with the reindex API. Plan the mapping for anything you will keep, and let dynamic mapping carry only throwaway experiments.

Mapping errors have a distinctive look. Index a document that contradicts the mapping:

HTTP
POST app-logs/_doc
{ "@timestamp": "2026-09-30T11:00:00Z", "host": "web-01", "level": "info", "status": "not-a-number", "message": "test" }

You get an HTTP 400 and a mapper_parsing_exception with the reason failed to parse field [status] of type [integer]. Read that message left to right: the exception class says the problem is the mapping, the field name tells you where, and the type tells you what was expected. The fix is either to correct the producer, to add an ingest conversion, or to map the field with "ignore_malformed": true so the bad value is dropped from the index while the rest of the document survives.

Another variation of the same idea is index templates, which apply a mapping and settings automatically to any new index or data stream whose name matches a pattern. That is how a fleet of daily or rolling indices all get the same schema without anyone creating them by hand. Templates are a central Mid-level topic. For now, know that when Filebeat or Elastic Agent writes a data stream, a template that the tools installed is what gave it its mapping.

Try it
  1. Create app-logs with the explicit mapping above.
  2. Index a valid document, then the one with "status": "not-a-number".
  3. Read the error and identify the exception name, the field, and the expected type.
the second request fails with a 400 and a mapper_parsing_exception. Finding those three pieces in the message is the skill; the error is telling you everything.

Aggregations: counting and grouping

Search answers "which documents match". Aggregations answer "what does the set of matching documents look like": how many per host, what is the average latency, how many errors per minute. They are what make dashboards possible, and Kibana generates them for you, but you should be able to read and write the basic ones.

An aggregation request goes in the same _search call, under aggs. Setting "size": 0 tells Elasticsearch you do not want the matching documents themselves, only the summary. On web-logs, where host and level have .keyword sub-fields, count documents per level:

HTTP
GET web-logs/_search
{
  "size": 0,
  "aggs": {
    "by_level": { "terms": { "field": "level.keyword" } }
  }
}

The response has aggregations.by_level.buckets, a list like { "key": "error", "doc_count": 2 }. The word bucket is the central noun: a bucket aggregation sorts documents into groups. The common bucket types are terms (one bucket per distinct value), date_histogram (one bucket per time interval) and range. Metric aggregations compute a number from the documents in a bucket: avg, sum, min, max, cardinality (count of distinct values) and percentiles. They nest, which is the powerful part:

HTTP
GET web-logs/_search
{
  "size": 0,
  "query": { "range": { "@timestamp": { "gte": "2026-09-30T00:00:00Z" } } },
  "aggs": {
    "per_minute": {
      "date_histogram": { "field": "@timestamp", "calendar_interval": "1m" },
      "aggs": {
        "errors": { "filter": { "term": { "level.keyword": "error" } } },
        "distinct_hosts": { "cardinality": { "field": "host.keyword" } }
      }
    }
  }
}

Read it from the outside in. The query narrows to one day. The date_histogram makes one bucket per minute. Inside every minute bucket, two sub-aggregations run: how many documents were errors, and how many distinct hosts appeared. That single request is the data behind a typical "errors over time" chart.

The most common aggregation error is this one: Fielddata is disabled on [host] in [web-logs]. Text fields are not optimised for operations that require per-document field data like aggregations and sorting. It appears when you aggregate on a text field. The remedy is in the message: aggregate on the keyword version, host.keyword, or map the field as keyword from the start. A second frequent one is too_many_buckets_exception, which appears when a request would create more than 65,536 buckets, for example a one-second histogram over a year. Use a larger interval or a narrower time range.

Try it
  1. Count documents per host.keyword with a terms aggregation and "size": 0.
  2. Add an avg of status inside each host bucket.
  3. Deliberately aggregate on host (without .keyword) and read the error.
three buckets for three hosts, an average per host, and the Fielddata error pointing you to the keyword field.

Kibana: seeing your data

Everything you did in Dev Tools is the raw interface. Kibana is how most people, including non-engineers, actually use the stack. Three places matter for a beginner: Discover, Dashboards, and Dev Tools (which you already know).

Data views and Discover

Kibana cannot browse an index until you tell it which indices to treat as one dataset. That object is a data view. Open the main menu, then Stack Management, then Data Views, and create one: give it a name like web-logs, type web-logs as the pattern (wildcards such as logs-* work), and choose @timestamp as the time field. The time field is what drives the time picker at the top of every Kibana page.

Now open Discover. It shows a histogram of documents over time, a table of documents below, and a search bar. Three habits make Discover productive:

  1. Set the time range first. The top-right picker defaults to a recent window. Your sample documents are from 30 September 2026, so widen the range (for example, "Last 2 years", or absolute dates) or you will see "No results found".
  2. Use the search bar with KQL, the Kibana Query Language. level : "error" and host : "web-01" is a valid query. KQL is simpler than the JSON Query DSL and is translated into it for you.
  3. Add fields as columns. Hover over a field in the left panel and click add, so the table shows host, level and message instead of one wide JSON blob.

Click any document to expand it and see every field and its mapped type. If a field you expected is missing, the usual cause is that the data view was created before the field existed; refresh the data view's field list in Stack Management.

Building a dashboard

A dashboard is a page of charts on the same time range. The fastest way to build your first is with Lens, Kibana's drag-and-drop chart editor. From the menu choose Dashboards, then Create dashboard, then Create visualization. Pick your data view, drag @timestamp to the horizontal axis and let Lens choose a count of records for the vertical axis: that is an "events over time" chart. Change it to a stacked bar and drag level into "break down by", and you have errors versus warnings over time. Save it, add it to the dashboard, add a second chart for top hosts, and save the dashboard with a name.

The mapping you chose pays off here. Lens only offers fields that can be aggregated for a given chart type, so keyword and numeric fields appear while text fields do not. If a field you want is missing from a chart's field picker, the cause is almost always its type.

ES|QL, the pipe language

During the 8.x series Elasticsearch gained a second query language that reads more like a shell pipeline: ES|QL. It is available in Discover and over the _query API, and on the free tier. Each command feeds the next:

ESQL
FROM web-logs
| WHERE level.keyword == "error"
| STATS errors = COUNT(*) BY host.keyword
| SORT errors DESC
| LIMIT 10

Read it as a sentence: from the index, keep errors, count them per host, sort the biggest first, keep ten. The equivalent in the Query DSL plus aggregations would be a page of JSON. Over the REST API, send it as POST /_query?format=txt with body { "query": "FROM web-logs | LIMIT 10" }. Two limits to remember: by default a query returns 1,000 rows, with a maximum of 10,000, and since version 9.1 ES|QL may return partial results when some shards fail, flagged by an is_partial field in the response, so look for it if numbers seem low.

For a beginner, the recommendation is: use KQL in Discover for quick finding, Lens for charts, and ES|QL when you want to ask a slightly more complex question in one readable query. You do not need to master the JSON Query DSL to be effective, but you should be able to read it, since documentation and error messages use it.

Stuck with no results in Discover? Go in this order First the time picker, then the data view (is it the right one?), then the search bar (clear it), then the field types. Ninety percent of "Discover is empty" questions end at the first step.
Try it
  1. Create a data view for web-logs with @timestamp as the time field.
  2. In Discover, widen the time range and search level : "error".
  3. Build a Lens bar chart of count over time, broken down by level, and save it to a new dashboard.
  4. Run the ES|QL query above and compare it to your chart.
the same error counts per host in the query as the error segments in your chart. One question, three ways to ask it.

Getting real logs in with Filebeat

Hand-typed documents taught you the API. Real systems produce logs continuously, and something must read the files and send each new line to Elasticsearch. That is a shipper, and the simplest one to learn is Filebeat: a tiny program that watches files, remembers how far it has read, and sends new lines on. It keeps a record of its position, so if Filebeat or the network fails, it resumes where it stopped instead of losing or duplicating lines.

Install Filebeat from the same repository as the other products (sudo apt-get install filebeat), or download the archive for your platform from Elastic's downloads page. Its configuration lives in one file, filebeat.yml. A complete minimal configuration for an application that writes to /var/log/app/*.log:

filebeat.yml
filebeat.inputs:
  - type: filestream
    id: app-logs
    paths:
      - /var/log/app/*.log

output.elasticsearch:
  hosts: ["http://localhost:9200"]
  username: "elastic"
  password: "${ES_PASSWORD}"

setup.kibana:
  host: "http://localhost:5601"

This example targets the plain-HTTP start-local stack. Four things deserve explanation.

The input type is filestream. Older tutorials use type: log. That input is deprecated, and in Filebeat 9 it will not start unless you add allow_deprecated_use: true. Use filestream for anything new.

The id is mandatory and must be unique. Filebeat 9 refuses to start if two inputs share an ID, because it tracks file positions per input. Give every input a meaningful, permanent name.

Credentials use a variable. ${ES_PASSWORD} is read from the environment, so the secret is not in the file. On a real server you would use Filebeat's own keystore or an API key instead of a user password.

Against a secured cluster, the output needs the certificate. For an HTTPS cluster the output becomes:

filebeat.yml
output.elasticsearch:
  hosts: ["https://es-01:9200"]
  api_key: "id:api_key"
  ssl.certificate_authorities: ["/etc/filebeat/http_ca.crt"]

Without the CA line, Filebeat logs x509: certificate signed by unknown authority, which is the shipper's version of curl's self-signed certificate complaint, and the fix is the same: provide the certificate rather than disabling verification.

Now the commands, in the order you should use them:

BASH
filebeat test config -e
filebeat test output
sudo filebeat setup -e
sudo systemctl start filebeat

test config checks the YAML for mistakes; test output proves Filebeat can reach and authenticate to Elasticsearch; setup loads the index template and sample dashboards into Elasticsearch and Kibana once; and start begins shipping. If test output succeeds and nothing arrives, check the file's paths and permissions next.

Filebeat also has modules that know how to read common software. filebeat modules enable nginx turns on a module that reads the nginx access and error logs, parses each line into named fields, and brings matching dashboards. Modules are what make "point it at nginx and get a dashboard" work with no parsing effort from you.

One Filebeat 9 behaviour that genuinely confuses people: the filestream input identifies files by a fingerprint of their first 1,024 bytes by default, so a file smaller than 1,024 bytes is not ingested until it grows. If you test with a tiny three-line file and nothing shows up, this is why. Append more text to the file, or write a longer test line.

Once data flows, open Discover. Filebeat writes to a data stream with a name like filebeat-9.5.4, or logs-* when you use Elastic Agent integrations. Create or use a data view matching it, and you will see your log lines with message holding the original text, plus fields describing the host and file.

Elastic Agent is the long-term path Elastic recommends Elastic Agent with Fleet for new deployments: one agent per machine, configured centrally from Kibana, with integrations that replace individual Beats and modules. You install it with sudo ./elastic-agent install --url=https://fleet-server:8220 --enrollment-token=<token> and check it with sudo elastic-agent status. The concepts (input, parsing, output) are identical to Filebeat, so learning Filebeat first costs you nothing. The Mid-level guide covers Fleet.
Try it
  1. Create a folder /tmp/demo-logs and a file in it with at least ten lines of text, so it is larger than 1,024 bytes.
  2. Write a filebeat.yml with one filestream input with an id, pointing at the file, and outputting to your local Elasticsearch.
  3. Run filebeat test config -e, then filebeat test output, then start Filebeat in the foreground with filebeat -e.
  4. Find the lines in Discover.
each line of your file as a document with a message field. If you see nothing, re-read the fingerprint paragraph.

Logstash: parsing and reshaping

A log line is text. 203.0.113.7 GET /checkout 500 123ms is readable by a person, but searchable structure (a client address, a method, a path, a status number, a duration number) only exists if something pulls those pieces apart. Filebeat modules do that for well-known software. For your own applications, you either parse at the source, or you use a tool built for it. Logstash is that tool: a pipeline that reads events, transforms them, and writes them out.

A Logstash pipeline has three stages, each written as a block in a .conf file:

INPUTwhere events come from
→
FILTERparse, enrich, clean
→
OUTPUTwhere events go

Inputs include files, TCP and HTTP listeners, message queues such as Kafka, and the beats input that receives from Filebeat on port 5044. Filters include grok (pattern-match text into fields), dissect (split by a fixed layout), date (set the event time from a field), mutate (rename, convert, drop), and geoip (turn an address into a location). Outputs include Elasticsearch, files, and the terminal.

Here is a complete pipeline for the log line shown above. Save it as pipeline.conf:

pipeline.conf
input {
  stdin { }
}

filter {
  grok {
    match => { "message" => "%{IP:client_ip} %{WORD:method} %{URIPATH:path} %{NUMBER:status:int} %{NUMBER:duration_ms:int}ms" }
  }
}

output {
  stdout { codec => rubydebug }
  elasticsearch {
    hosts    => ["http://localhost:9200"]
    user     => "elastic"
    password => "${ES_PASSWORD}"
    index    => "shop-logs"
  }
}

Every line of it has a job. The stdin input reads whatever you type, which makes it a perfect learning tool. The grok filter holds a pattern made of named building blocks: %{IP:client_ip} means "match something shaped like an IP address and store it in the field client_ip", and :int converts the captured text into a number so that it can be averaged later. The stdout output with the rubydebug codec prints each finished event in a readable structure, which is how you debug a pipeline. The elasticsearch output indexes into shop-logs.

Before running anything, test the syntax. Logstash's startup is slow, and a typo found after thirty seconds is a waste:

BASH
bin/logstash -f pipeline.conf --config.test_and_exit

Then run it, type a line, and watch:

BASH
export ES_PASSWORD="<the password from .env>"
bin/logstash -f pipeline.conf
# now type:
203.0.113.7 GET /checkout 500 123ms

The debug output shows the event with client_ip, method, path, status and duration_ms as separate fields, plus @timestamp (set to the time Logstash received it), the original message, and a host or similar field describing where it ran. If instead you see "tags" => ["_grokparsefailure"], the pattern did not match the line. That tag is Logstash's way of saying "I could not parse this", and it is the single most common Logstash symptom. Fix it by comparing your pattern against the line character by character; the Grok Debugger in Kibana Dev Tools lets you paste a line and a pattern and see what matches.

The Elasticsearch output keeps the secret out of the file too: "${ES_PASSWORD}" is expanded from the environment, and bin/logstash-keystore can hold it on a server. When you move to HTTPS, the option names in version 9 matter: use ssl_certificate_authorities for the CA file, not the old cacert. Logstash 9's Elasticsearch output refuses to start if the removed names (cacert, ssl, keystore) are present. For a real cluster, use an API key with api_key => "id:api_key", scoped to only the indices the pipeline writes.

Two more facts about running Logstash. It needs a newer Java than older tutorials assume, but it bundles its own, so ignore that; and as of version 9 it refuses to run as root (allow_superuser defaults to false). Run it as the logstash user that the package creates. And if you start a second Logstash with the same data directory you get Logstash could not be started because there is already another instance using the configured data directory, so use one instance per data path.

When do you actually need Logstash? Not always. Use it when you need heavy or conditional parsing, enrichment from another source, routing one stream to several destinations, or a buffer in front of Elasticsearch. Skip it when Filebeat modules or Elastic Agent integrations already parse your source, because every extra component is one more thing to run and to monitor.

Try it
  1. Save the pipeline above and run the syntax test.
  2. Start Logstash and paste three lines of the same shape with different status codes and durations.
  3. Paste one deliberately wrong line, such as hello world, and find the _grokparsefailure tag.
  4. In Discover, create a data view for shop-logs and chart the average duration_ms.
three parsed events, one tagged failure, and a numeric chart, which only works because of the :int conversion.

Configuration you will actually touch

Three products, three configuration files, and a short list of settings that cover your first months. Each file is YAML or a close relative, and each product reads secrets from a keystore or environment variable rather than the file when you ask it to.

elasticsearch.yml

On a package install it lives in /etc/elasticsearch/. A single-node development file needs very little:

elasticsearch.yml
cluster.name: learning
node.name: es-01
path.data: /var/lib/elasticsearch
path.logs: /var/log/elasticsearch
network.host: 127.0.0.1
http.port: 9200
discovery.type: single-node

cluster.name and node.name are labels; give them meaningful values because they show up in every log and health check. path.data is where the indexed data lives, which should be outside the install folder so an upgrade never touches it. discovery.type: single-node tells Elasticsearch not to look for other nodes; it is for development only.

A critical behaviour: binding to a non-loopback address changes the rules. While network.host stays on localhost, Elasticsearch runs in development mode and only warns about poor settings. The moment you set it to a real address such as 0.0.0.0, it switches to production mode and turns those warnings into bootstrap checks: hard failures that stop the node from starting. That is when beginners meet max virtual memory areas vm.max_map_count [65530] is too low and similar. Those are not bugs; they are Elasticsearch refusing to run in a configuration that is likely to lose data.

For a multi-node cluster you replace discovery.type with discovery.seed_hosts (the list of nodes to contact) and, for the very first start only, cluster.initial_master_nodes. Remove initial_master_nodes afterwards, and never add it to a node joining an existing cluster.

Memory is the other setting people expect to touch, and usually should not. The heap is sized automatically from the node's RAM and roles. If you must override it, put -Xms and -Xmx (equal values) in a file under jvm.options.d/, never in jvm.options itself, and keep the heap at no more than half of the machine's memory.

kibana.yml

kibana.yml
server.host: "localhost"
server.port: 5601
elasticsearch.hosts: ["https://localhost:9200"]
elasticsearch.ssl.certificateAuthorities: ["/etc/kibana/certs/http_ca.crt"]

Kibana needs to know where Elasticsearch is, and, for HTTPS, which certificate authority to trust. It authenticates as a built-in service identity (the enrollment flow sets that up), never as the elastic superuser; Kibana refuses elasticsearch.username: elastic. To let other machines reach Kibana, change server.host to 0.0.0.0, and put it behind HTTPS before you do.

If you use alerting or Fleet you will later be asked for encryption keys, xpack.encryptedSavedObjects.encryptionKey among them, at least 32 characters. Generate them with bin/kibana-encryption-keys generate, store them somewhere safe, and use the same values on every Kibana instance. Lose them, and the saved data they encrypt becomes unreadable.

logstash.yml and pipelines.yml

logstash.yml holds process-level settings; defaults are fine to start. The ones you will meet are pipeline.workers (defaults to your CPU count), pipeline.batch.size (125), queue.type (memory by default; persisted writes events to disk so a crash does not lose them), and config.reload.automatic, which makes Logstash pick up pipeline edits without a restart, also available as --config.reload.automatic on the command line. The file pipelines.yml lists several independent pipelines, each with a pipeline.id and a path.config.

Secrets

The habit to adopt on day one: no passwords in config files that go into git. Elasticsearch, Kibana and Logstash each have a keystore command (elasticsearch-keystore, kibana-keystore, logstash-keystore), and all three configs expand ${VARIABLE} from the environment. Create API keys for machines rather than sharing the elastic password. An API key can be limited to certain indices and given an expiry, so a leaked key is a bounded problem:

HTTP
POST /_security/api_key
{
  "name": "filebeat-writer",
  "expiration": "90d",
  "role_descriptors": {
    "writer": {
      "cluster": ["monitor"],
      "indices": [
        { "names": ["logs-*", "filebeat-*"], "privileges": ["auto_configure", "create_doc"] }
      ]
    }
  }
}

The response contains id, api_key and encoded. Beats and Logstash take the pair as id:api_key; raw HTTP clients send the encoded value in an Authorization: ApiKey <encoded> header. Copy the key when it is displayed; Elasticsearch cannot show it again. Note that elasticsearch-setup-passwords is deprecated; use elasticsearch-reset-password when you need a new password for a built-in user.

Try it
  1. Find your elasticsearch.yml or, with start-local, the settings in the Compose file, and locate the cluster name.
  2. Create an API key with the request above and use it in curl -H "Authorization: ApiKey <encoded>" http://localhost:9200.
  3. Invalidate it with DELETE /_security/api_key and the returned id in ids.
the key works, then stops working after deletion, with a 401. That is how machine credentials should behave: easy to create, easy to revoke.

Common errors and how to read them

Elasticsearch errors are long, but they are well structured. Nearly all follow the same layout: an HTTP status (400 your request is wrong, 401 or 403 security, 404 not found, 429 the cluster is pushing back), an exception type, and a reason sentence. Read the type first, then the reason; it often contains the fix.

The node will not start

max virtual memory areas vm.max_map_count [65530] is too low, increase to at least [262144]. A bootstrap check in production mode. Elasticsearch memory-maps its files and needs many mappings. Fix it on the host: sudo sysctl -w vm.max_map_count=1048576 (the value the 9.x documentation recommends; the check only requires 262144), and persist it in /etc/sysctl.conf. Package installs set this automatically. With Docker Desktop, it must be set inside the VM Docker uses.

max file descriptors [4096] for elasticsearch process is too low, increase to at least [65535]. Raise the limit with ulimit -n 65535, or LimitNOFILE=65535 in the systemd unit.

the default discovery settings are unsuitable for production use; at least one of [discovery.seed_hosts, discovery.seed_providers, cluster.initial_master_nodes] must be configured. You bound to a real address without telling the node how to find a cluster. Set discovery.type: single-node for one node, or configure the discovery settings for several.

initial heap size [X] not equal to maximum heap size [Y]. Set -Xms and -Xmx to the same value in jvm.options.d/.

Clients cannot connect or authenticate

received plaintext http traffic on an https channel, closing connection. Elasticsearch's log, after a client used http:// against the TLS port. Use https://.

SSL certificate problem: self-signed certificate in certificate chain from curl, or x509: certificate signed by unknown authority from Filebeat. The client does not trust the cluster's certificate authority. Pass --cacert http_ca.crt to curl, or set ssl.certificate_authorities in Beats.

HTTP 401, missing authentication credentials for REST request or unable to authenticate user. No credentials, or wrong ones. Reset the password with bin/elasticsearch-reset-password -u elastic.

HTTP 403, action [indices:data/write/bulk[s]] is unauthorized for user [x]. You authenticated, but your role lacks the privilege. For data streams, give the writer create_doc and auto_configure on the index pattern.

Kibana: Kibana server is not ready yet. Either Kibana is still starting (wait a minute) or it cannot talk to Elasticsearch. Check that elasticsearch.hosts uses the right scheme, that the CA is trusted, that Elasticsearch is not red, and that both run the same version. The Kibana log tells you which.

Data problems

mapper_parsing_exception … failed to parse field [status] of type [integer]. A value contradicts the mapping. Covered above.

cluster_block_exception … disk usage exceeded flood-stage watermark, index has read-only-allow-delete block. The disk is 95% full, and Elasticsearch protects itself by making indices read-only. Free space or add capacity. The block lifts on its own when usage falls below the high watermark.

circuit_breaking_exception: [parent] Data too large. HTTP 429. A request would use too much heap. Make the query or aggregation smaller, and never aggregate on text fields.

Limit of total fields [1000] has been exceeded. An index has more than 1,000 fields, almost always because dynamic mapping made a field per key of some free-form object. Switch that object to flattened, set dynamic to strict or false, or map explicit fields.

Result window is too large, from + size must be less than or equal to: [10000]. You paged too deep. Use narrower queries now; search_after later.

Shipper problems

Filebeat exits because of a duplicate input ID, or the log input is deprecated. Give each filestream input a unique id, and stop using type: log.

Exiting: data path already locked by another beat. Two Filebeat processes share one data folder. Stop one.

Logstash _grokparsefailure. Your pattern does not match the line. Test it in the Grok Debugger.

Do not silence security errors Searching for an error will turn up advice to add -k to curl, set ssl.verification_mode: none, or turn xpack.security.enabled off. Each makes the message vanish and each removes a protection. Do that in a throwaway lab at most, and never write it down as the solution.
Try it
  1. Run curl https://localhost:9200 against a TLS-enabled node without --cacert and note the error.
  2. Send a request to a wrong index name, for example GET nope/_search, and read the exception type.
  3. Index a document with a wrong-typed field into app-logs and identify the field in the message.
three errors, each read as type, then reason, then fix. That reading order is the whole skill.

Putting it all together

Here is a small project that uses everything above. You will build a tiny log-monitoring setup for a fictional model-serving application, from file to dashboard, in under an hour.

  1. Start the stackRun the start-local script, read the password from .env, and confirm GET _cluster/health in Dev Tools.
  2. Design the mappingCreate a model-logs index with @timestamp as date, host and level as keyword, latency_ms as integer, status as integer, and message as text.
  3. Generate a log fileWrite a short script that appends a line such as 2026-09-30T10:00:05Z web-01 INFO 200 34ms predict ok every second with random hosts, levels, statuses and latencies, to a file, making sure it passes 1,024 bytes.
  4. Parse with LogstashWrite a pipeline with a file or stdin input, a grok filter that extracts timestamp, host, level, status and latency as numbers, and a date filter that uses the timestamp as @timestamp. Output to model-logs. Run it with --config.test_and_exit first.
  5. Check the dataUse GET model-logs/_count and a terms aggregation on level to confirm the numbers look right, and search for one specific error message.
  6. Explore in KibanaCreate a data view, find the errors in Discover with KQL, then build a dashboard with count over time by level, average latency per host, and a table of the ten latest errors.
  7. Break it on purposeAdd a line with latency_ms that is not a number and watch both the _grokparsefailure tag and, if you bypass grok, the mapper error.
  8. Clean upDelete the index with DELETE model-logs and stop the stack with ./stop.sh.

The result is an end-to-end pipeline you understand rather than one you copied. You can name every stage, explain what each one did, and recognise the typical failure at each step.

What you can now do, and what comes next

You can start and verify a local stack, put data in by hand, in bulk and through shippers, search with match, term, range and bool, summarise with aggregations and ES|QL, set explicit mappings, parse text with Logstash, build a Kibana dashboard, and read the most common errors. That is enough to be useful on a team that already runs the stack.

What this guide deliberately left out: how indices age and get deleted automatically (index lifecycle management and data stream lifecycle), index and component templates in depth, ingest pipelines that parse inside Elasticsearch without Logstash, Fleet and Elastic Agent management, roles and users for a team, multi-node clusters, and upgrades. All of that is the Mid-level guide. Running the stack as a platform for other teams (sizing, security review, failure handling) is the Senior guide.

Natural neighbours in this catalogue: containers and Docker for running the stack itself (Docker), and Kubernetes for running it at scale (Kubernetes). When you deploy model services, the logs they produce are what this stack is for.

Before you go, re-run the Sixty-second self-test in the interview guide and skim the tips guide's checklist. Explaining the stack out loud is the quickest way to find what you only half understood.

Sources