Skip to content
Back to student guides
QdrantLLMsVector databases3 levels101 sectionsCovers Qdrant 1.19

The Complete Qdrant Guide

Run fast, filterable vector search with the open-source Qdrant engine. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
19sections
28examples

This is part one of three. It covers everything you need to do real work with Qdrant on your first day and through your first month. By the end you can start a Qdrant server, create a collection, store vectors with metadata, search by meaning, narrow results with filters, read the common errors, and keep the whole thing from leaking onto the public internet. Mid-level and Senior take the same ground further (hybrid search, quantization, clustering, operations); nothing here is thrown away.

The guide targets Qdrant v1.19.1, the stable release at the time of writing. Qdrant moves quickly, so some older tutorials on the internet show endpoints and settings that have since been deprecated. Where that matters, this guide says so and shows the current form.

Each section ends with a Try it task. Do them as you go. Vector search feels abstract until you have watched your own six points come back in a sensible order.

The problem Qdrant solves

A normal database answers exact questions. "Give me the order with id 4021." "Give me every customer where country is Egypt." The database compares values for equality or order, and a row either matches or it does not.

A lot of modern software needs a different kind of question: "give me the things most like this one." Find the support articles closest in meaning to a customer's complaint. Find the product photos that look like this photo. Find the paragraphs in our documents that are relevant to a question. None of these have an exact answer. They have a ranking.

The trick that makes this possible is the embedding. A machine learning model reads a piece of content (a sentence, an image, an audio clip) and produces a list of numbers, typically a few hundred to a few thousand of them. That list is a vector. The model is trained so that content with similar meaning lands at nearby positions in that number space. "How do I reset my password?" and "I forgot my login" end up close together, even though they share almost no words. "Quarterly revenue report" ends up far away.

Once everything is a vector, "most like this" becomes a geometry problem: given a query vector, find the stored vectors with the smallest distance to it. That is called nearest-neighbour search, and it is exactly what Qdrant does.

Qdrant (pronounced "quadrant") is an open-source vector database and vector search engine, licensed under Apache 2.0 and written in Rust. You talk to it over a REST API on port 6333 or a gRPC API on port 6334. It stores your vectors, stores a JSON object of metadata next to each one, builds an index so searches stay fast when you have millions of vectors, and lets you combine similarity with ordinary filters such as "only documents from 2025" or "only this customer's data".

CONTENTtext, image, audio
→
EMBEDDING MODELcontent to numbers
→
QDRANTstore + index
→
NEAREST POINTSranked by similarity

The diagram is the whole idea. Notice what Qdrant is not. It does not create embeddings for you when you run it yourself; the embedding model is a separate step that runs before data reaches Qdrant. (Qdrant offers helper libraries and, on its managed cloud, server-side embedding, but the mental model stays the same: vectors in, nearest vectors out.) It is also not a replacement for your main application database. Most teams keep orders and users in a relational database and keep only the searchable representation, plus enough metadata to link back, in Qdrant.

Try it
  1. Write down three search features you have used this week (a store search box, a music "similar songs" list, a photo app that finds "beach").
  2. For each one, decide whether the question was exact ("order 4021") or "most like this".
  3. Pick one "most like this" feature and name what content would have to be turned into vectors to build it.

What came before, and why a dedicated engine

You could do nearest-neighbour search without Qdrant. The naive method is to keep all vectors in a list and compare the query against every one of them. That is called a brute-force or exact scan. It gives perfect answers, and for a few thousand vectors it is perfectly fine. The cost grows linearly, though. With ten million vectors of 768 numbers each, every single query would do billions of multiplications, which is far too slow for an interactive application.

The classic fix is an approximate nearest-neighbour index. Instead of guaranteeing the exact closest vectors, it finds vectors that are almost certainly among the closest, in a tiny fraction of the time. You trade a sliver of accuracy for speed that is orders of magnitude better. Libraries such as FAISS (see the FAISS guide) implement these indexes, but a library is only the index. You still have to build storage, persistence, crash recovery, updates, deletes, metadata filtering, an API, and replication around it.

A vector database is that whole package. Other options in the same family are covered elsewhere in this catalogue: Pinecone (managed service), Weaviate, Milvus, Chroma, and pgvector (vector search inside PostgreSQL). Qdrant's particular strengths are a clean filtering model, a single flexible query API, and a Rust engine you can run as one container on a laptop and later as a cluster.

Two practical consequences shape how you use it. First, search is approximate by default. Two runs over the same data can differ at the margin, and a result you expect may occasionally be missing at the very bottom of a list. For nearly every real use that is acceptable, and you can ask for an exact search when you need one. Second, every vector in a given space must have the same length. You cannot mix 384-number vectors and 768-number vectors in the same space, because distance between them is meaningless. This catches almost every beginner once.

Pick the embedding model first The model decides the vector length and which distance metric suits it. Decide that before you create a collection, because the length is fixed when you create it. If you switch models later, you generate every vector again and load them into a new collection.
Try it
  1. Estimate the cost of brute force: multiply 10,000,000 vectors by 768 numbers. That is the number of multiplications per query.
  2. Compare it with a budget of 50 milliseconds and decide why an index is needed.

The mental model: collections, points, vectors, payloads

Qdrant has four nouns. Learn them precisely, because every API call is built from them.

A collection is a named set of points that you search together. It is the nearest thing to a table. You might have a collection called support_articles and another called product_images. When you create a collection you fix the length of its vectors and the distance metric it uses. Search always happens inside one collection.

A point is the central record. Every point has an ID, zero or more vectors, and an optional payload. The ID is either an unsigned 64-bit integer or a UUID string. IDs are how you update, fetch, and delete a specific point; if you upsert a point with an ID that already exists, you replace the old one.

A vector is the list of floating-point numbers. In the simplest setup each point has exactly one. Qdrant also supports several named vectors per point (for example an image vector and a text vector, each with its own length), and sparse vectors that store only the non-zero entries for keyword-style matching. Beginners start with one unnamed dense vector and add the rest later.

A payload is a JSON object attached to a point. It holds the metadata: the original text, a title, a URL, a category, a date, a price, a tenant ID, anything you want to return with results or filter by. Payload values can be strings, numbers, booleans, geo-coordinates, dates, and nested objects or arrays.

Collection: cities (4 numbers per vector, Cosine)
Point id 1
vector [0.1, 0.9, 0.7, 0.1]
payload {"city": "Cairo", "country": "Egypt"}
Point id 2
vector [0.9, 0.7, 0.4, 0.4]
payload {"city": "Alexandria", "country": "Egypt"}
Point id 3
vector [0.7, 0.2, 0.9, 0.3]
payload {"city": "Dubai", "country": "UAE"}

A few extra terms appear in the docs and in error messages, so meet them now.

  • Segment. Inside a collection, points are stored in chunks called segments. Each segment has its own vector storage, payload storage, and indexes. You rarely touch them directly, but they explain why a brand-new collection behaves differently from a large one.
  • WAL (write-ahead log). Every change is written to a log on disk before it is applied. Once Qdrant has acknowledged a write, it survives a power cut.
  • Optimizer. A background process that rebuilds segments, merges small ones, cleans up deleted points, and builds indexes. It is why a collection can show a "yellow" status for a while after a big upload.
  • Shard and replica. A collection can be split into shards, and shards can be copied as replicas on other servers. On one machine, the default is simply to work; you will care about these when you run a cluster.
  • Alias. An extra name for a collection, which lets you swap one collection for another without changing application code. A mid-level topic.
Think of it as a table with a magnet A collection is like a table where each row (point) has an ID, a JSON column (payload), and a special column (vector) that you can query by "closest to this". Filters work on the JSON column; similarity works on the vector column; the interesting queries use both.
Try it
  1. Sketch a collection for a company knowledge base. Decide what the point ID is, what text becomes a vector, and which three payload fields you would store.
  2. Say which of those payload fields you would filter on, and which you would only display.

Distance metrics and how to read a score

When you create a collection you choose how "closeness" is measured. Qdrant offers four metrics.

  • Cosine compares the direction of two vectors and ignores their length. Two vectors pointing the same way score as identical even if one is much longer. Qdrant implements it by normalising each vector when you upload it and then using a dot product. Most text-embedding models are designed for cosine similarity, so it is the usual default.
  • Dot is the dot product, which depends on both direction and length. Use it when the model's documentation says to, which includes some recommendation models and sparse vectors.
  • Euclid is straight-line distance between the points.
  • Manhattan is the sum of absolute differences along each dimension.

The most useful rule is: use the metric your embedding model was trained for. The model's documentation or model card states it. Picking a different one does not crash anything; it silently gives worse rankings.

Reading a score needs one more piece of care, because the direction differs. For Cosine and Dot, a higher score means closer. For Euclid and Manhattan the value is a distance, so a lower number means closer. This matters when you use score_threshold, a search option that discards results scoring worse than a cut-off. A threshold of 0.8 on a Cosine collection keeps the good matches; the same number on a Euclid collection means something quite different.

Scores are not percentages A Cosine score of 0.9 does not mean "90 percent relevant". Scores only rank results against each other for one query, one model, one collection. Do not hard-code a global cut-off such as "reject anything under 0.75" without measuring what your own data produces.
Try it
  1. Find the model card of an embedding model you might use and write down its vector length and recommended similarity measure.
  2. Decide which of Cosine, Dot, Euclid, or Manhattan that corresponds to.

How Qdrant finds neighbours quickly

You can use Qdrant for a long time without knowing how the index works, but a little intuition prevents confusion about settings and about why results are "approximate".

Qdrant's index for dense vectors is HNSW, short for Hierarchical Navigable Small World. Picture a road network. To get from one city to another you do not check every road. You take a highway to get roughly close, then a regional road, then local streets. HNSW builds a layered graph over your vectors in the same spirit. The top layer has few points and long links; the lower layers have more points and shorter links. A search enters at the top, greedily hops toward the query, drops down a layer, and repeats until it reaches the bottom, where it collects the closest candidates.

Two build-time settings control the graph. m is the number of links per point (default 16); more links means better recall and more memory. ef_construct is how many candidates are considered while building (default 100); higher builds a better graph more slowly. At search time, hnsw_ef is how many candidates the search keeps track of; higher is more accurate and slower. The defaults are sensible. As a beginner, resist the urge to tune them until you can measure a problem.

Two behaviours follow from this design and surprise newcomers.

Small collections are not indexed at all. Qdrant does not build an HNSW graph for a segment until it passes an indexing threshold (10,000 KB of vector data by default, which is a few thousand vectors of typical size). Below that, a plain scan is quicker than a graph, so Qdrant just scans. If you create a six-point collection and see indexed_vectors_count of 0, nothing is broken.

Filtering and the index interact. Qdrant can combine a filter with the graph search, but it does that best when the filtered fields have a payload index. We return to this in the filtering sections, because it is the single most valuable habit to learn early.

Approximate does not mean sloppy In practice HNSW returns the true nearest neighbours for the vast majority of queries. When you need certainty (for example to measure how good the approximation is), a search can request an exact scan instead. Treat exact mode as a testing tool, not the default.
Try it
  1. Explain in two sentences, to a friend, why asking a map app for a route does not check every road in the country.
  2. Then say which part of HNSW corresponds to taking the highway.

Installing Qdrant

The fastest way to run Qdrant for learning is Docker. If Docker is new to you, read the Docker guide first; you only need to run one container, but it helps to know what a container and a volume are.

Pull the image and start the server, pinning the version so your results match this guide:

BASH
docker pull qdrant/qdrant:v1.19.1
docker run -p 6333:6333 -p 6334:6334 \
    -v "$(pwd)/qdrant_storage:/qdrant/storage:z" \
    qdrant/qdrant:v1.19.1

Read the command in pieces. -p 6333:6333 publishes the REST API and the Web UI; -p 6334:6334 publishes gRPC. The -v flag mounts a folder from your machine to /qdrant/storage inside the container, which is where Qdrant keeps all its data. Without that mount, your collections would vanish the moment the container is removed. The :z suffix asks SELinux-enabled Linux systems to relabel the folder so the container can write to it; it is harmless elsewhere. Snapshots, a backup mechanism, go to /qdrant/snapshots inside the container.

You will see log lines ending with the server listening on its ports. Leave this terminal running and open a second one for the rest of the guide. Press Ctrl+C in the first to stop the server.

On Windows

If you use Docker Desktop on Windows with WSL, do not bind-mount a Windows folder into the container. That kind of mount is not fully POSIX-compatible, and Qdrant's storage engine depends on POSIX behaviour. The documented failures are serious: vectors zeroed after a restart and panics mentioning OutputTooSmall. Use a named volume, which Docker keeps on its own Linux filesystem:

BASH
docker volume create qdrant-storage
docker run --rm -it -p 6333:6333 -p 6334:6334 \
    -v qdrant-storage:/qdrant/storage \
    qdrant/qdrant:v1.19.1

Other ways to install

Qdrant also ships as downloadable binaries for Linux, macOS (Intel and Apple Silicon), and Windows from the GitHub releases page, as a .deb package, and as an AppImage. You can build it from source with cargo build --release --bin qdrant. On Kubernetes there is a community-supported Helm chart (helm repo add qdrant https://qdrant.to/helm, then helm install qdrant qdrant/qdrant). Only 64-bit x86 and ARM64 are supported; 32-bit machines are not.

Qdrant's own documentation recommends its managed Qdrant Cloud, or its Kubernetes operator, for production. If you run it yourself with Docker, you take on high availability, backups, security, and monitoring. For learning and for single-server projects, the container above is a good start. If your employer has data-residency rules, which is common for Gulf and Egyptian organisations handling customer data, self-hosting on a server in an approved region keeps the vectors and the payloads where policy requires.

The storage rule you must know

Qdrant's data directory must be on a POSIX-compatible block filesystem, ideally SSD or NVMe. Network filesystems such as NFS and object storage such as S3 are not supported for the storage path. Since version 1.15 Qdrant checks the filesystem at start-up and warns or refuses. On macOS you may see a warning like "HFS/HFS+ filesystem support is untested". That is a warning, not an error.

Do not run two servers on one data folder If you start a second container pointing at the same qdrant_storage folder you get Can't open Collections meta Wal ... Resource temporarily unavailable. One storage folder belongs to exactly one Qdrant process. Give every instance its own folder or volume.
Try it
  1. Start Qdrant with the command above.
  2. Stop it with Ctrl+C, start it again with the same command, and notice that nothing about the folder needs recreating.
  3. Look inside qdrant_storage on your machine and see the files Qdrant created.

Checking that the setup works

Before building anything, confirm the server is alive and which version you have. From a second terminal:

BASH
curl http://localhost:6333

The reply is a small JSON welcome message containing the title and the server version. If it says 1.19.1, you are on the version this guide describes.

BASH
curl http://localhost:6333/healthz

A healthy server answers with HTTP 200 and a short text. Two siblings exist: /livez and /readyz. Kubernetes uses them as liveness and readiness probes. They stay reachable even when you later turn on an API key, which is exactly what you want from a health check.

BASH
curl http://localhost:6333/collections

On a fresh node this returns an empty list of collections: {"result":{"collections":[]},"status":"ok", ...} plus a time field.

Now open http://localhost:6333/dashboard in your browser. This is the built-in Web UI. It lists collections, lets you inspect points, and includes a Console where you can type the same REST requests this guide uses and see responses, which is a good alternative to curl if you prefer a browser. There is also an interactive tutorial at /dashboard#/tutorial. A graphical view of the vectors helps build intuition once you have data.

Qdrant also exposes /metrics in Prometheus format. You do not need it today, but it is good to know it is there; monitoring is a mid-level topic.

Always add ?wait=true to writes from curl The Python, JavaScript, .NET, and Java clients wait for a write to be applied before returning. Raw REST requests, and the Go and Rust clients, do not by default. If you upsert with curl and immediately search, the new points may not be visible yet. Add ?wait=true to the URL and the request returns only after the change has been applied.
Try it
  1. Run the three curl commands and read each reply. Find the version number in the first.
  2. Open the dashboard and click through to the Console.
  3. Run GET collections from the Console and compare it with the curl output.

Your first project: six cities and four numbers

Real embeddings have hundreds of dimensions, which makes them impossible to read. To learn the API, we use an invented toy dataset where each vector has just four numbers that we can reason about. Pretend each city is scored from 0 to 1 on four traits, in this order: sea, history, nightlife, quiet. The scores are made up for teaching; they are not data about the cities.

ID City Vector (sea, history, nightlife, quiet) Country Budget
1 Cairo 0.1, 0.9, 0.7, 0.1 Egypt low
2 Alexandria 0.9, 0.7, 0.4, 0.4 Egypt low
3 Dubai 0.7, 0.2, 0.9, 0.3 UAE high
4 Muscat 0.8, 0.6, 0.3, 0.8 Oman mid
5 Casablanca 0.8, 0.5, 0.6, 0.4 Morocco mid
6 Amman 0.0, 0.8, 0.5, 0.6 Jordan mid

Step 1: create the collection

BASH
curl -X PUT 'http://localhost:6333/collections/cities' \
  -H 'Content-Type: application/json' \
  -d '{"vectors": {"size": 4, "distance": "Cosine"}}'

PUT /collections/{name} creates a collection. The body says every vector has 4 numbers and similarity is Cosine. You get {"result":true,"status":"ok",...}. Run it a second time and Qdrant refuses, because the collection already exists; creation is not silently repeated.

Step 2: insert the points

Inserting is called an upsert: insert if the ID is new, replace if it exists.

BASH
curl -X PUT 'http://localhost:6333/collections/cities/points?wait=true' \
  -H 'Content-Type: application/json' \
  -d '{
    "points": [
      {"id": 1, "vector": [0.1, 0.9, 0.7, 0.1], "payload": {"city": "Cairo", "country": "Egypt", "budget": "low"}},
      {"id": 2, "vector": [0.9, 0.7, 0.4, 0.4], "payload": {"city": "Alexandria", "country": "Egypt", "budget": "low"}},
      {"id": 3, "vector": [0.7, 0.2, 0.9, 0.3], "payload": {"city": "Dubai", "country": "UAE", "budget": "high"}},
      {"id": 4, "vector": [0.8, 0.6, 0.3, 0.8], "payload": {"city": "Muscat", "country": "Oman", "budget": "mid"}},
      {"id": 5, "vector": [0.8, 0.5, 0.6, 0.4], "payload": {"city": "Casablanca", "country": "Morocco", "budget": "mid"}},
      {"id": 6, "vector": [0.0, 0.8, 0.5, 0.6], "payload": {"city": "Amman", "country": "Jordan", "budget": "mid"}}
    ]
  }'

Each point carries an ID, a vector, and a payload, exactly the three things from the mental model. The response reports a status such as completed for the operation.

Suppose a traveller wants somewhere seaside, with some history, a little nightlife, and fairly quiet. Express that as a query vector in the same four-number space, and ask for the three nearest points:

BASH
curl -X POST 'http://localhost:6333/collections/cities/points/query' \
  -H 'Content-Type: application/json' \
  -d '{"query": [0.9, 0.6, 0.3, 0.5], "limit": 3, "with_payload": true}'

POST /collections/{name}/points/query is the Query API, Qdrant's single search endpoint. The reply is abbreviated here:

TEXT
{"result": {"points": [
   {"id": 2, "score": 0.99, "payload": {"city": "Alexandria", ...}},
   {"id": 4, "score": 0.97, "payload": {"city": "Muscat", ...}},
   {"id": 5, "score": 0.96, "payload": {"city": "Casablanca", ...}}
 ]}, "status": "ok", ...}

Read it like this. The points are sorted best first. Each has its id, a score (Cosine, so higher is closer), and, because we asked for with_payload, its metadata. Without with_payload: true the payload is omitted; Qdrant returns only what you request, to save bandwidth. The same goes for vectors: add "with_vector": true to get them back. The scores you see will be close to the ones above; the order is what matters. Alexandria is the best match because its profile (strong sea, good history, calm-ish) is nearly the same direction as the query.

Step 4: add a filter

Now the traveller says: only mid-budget cities.

BASH
curl -X POST 'http://localhost:6333/collections/cities/points/query' \
  -H 'Content-Type: application/json' \
  -d '{
    "query": [0.9, 0.6, 0.3, 0.5],
    "filter": {"must": [{"key": "budget", "match": {"value": "mid"}}]},
    "limit": 3,
    "with_payload": true
  }'

Alexandria disappears, because its budget is low. The remaining mid-budget cities are ranked: Muscat, Casablanca, Amman. That is the core pattern of Qdrant in a single request: similarity decides the order, the filter decides who is allowed in.

Try it
  1. Run all four steps in order. After step 2, open the dashboard and view the cities collection.
  2. Change the query vector to [0.0, 0.9, 0.5, 0.3] (history-heavy, no sea) and predict the top result before running it.
  3. Filter on country equal to Egypt and confirm only Cairo and Alexandria can appear.

The same project in Python

The REST API is the foundation, but most people use a client library. The official Python client is qdrant-client:

BASH
pip install qdrant-client

If you later want Qdrant's local embedding library, FastEmbed, install the extra with pip install "qdrant-client[fastembed]". We do not need it for the toy vectors. Official clients also exist for JavaScript/TypeScript (@qdrant/js-client-rest), Rust, Go, .NET, and Java; they mirror the same REST and gRPC calls.

cities.py
from qdrant_client import QdrantClient
from qdrant_client.models import (
    Distance, VectorParams, PointStruct,
    Filter, FieldCondition, MatchValue,
)

client = QdrantClient(url="http://localhost:6333")

client.create_collection(
    collection_name="cities_py",
    vectors_config=VectorParams(size=4, distance=Distance.COSINE),
)

client.upsert(
    collection_name="cities_py",
    wait=True,
    points=[
        PointStruct(id=1, vector=[0.1, 0.9, 0.7, 0.1],
                    payload={"city": "Cairo", "country": "Egypt", "budget": "low"}),
        PointStruct(id=2, vector=[0.9, 0.7, 0.4, 0.4],
                    payload={"city": "Alexandria", "country": "Egypt", "budget": "low"}),
        PointStruct(id=4, vector=[0.8, 0.6, 0.3, 0.8],
                    payload={"city": "Muscat", "country": "Oman", "budget": "mid"}),
    ],
)

hits = client.query_points(
    collection_name="cities_py",
    query=[0.9, 0.6, 0.3, 0.5],
    query_filter=Filter(must=[FieldCondition(key="budget", match=MatchValue(value="mid"))]),
    limit=3,
).points

for hit in hits:
    print(hit.id, round(hit.score, 3), hit.payload["city"])

The structure mirrors the REST calls one to one. QdrantClient(url=...) connects; add api_key="..." when the server requires one. create_collection takes a VectorParams object holding the same size and distance. upsert takes PointStruct objects. query_points takes the query vector, an optional query_filter, and a limit, and its .points attribute holds the scored results. The Python client waits for writes by default, so wait=True is stating the default; keeping it is a good way to make intent obvious.

Do not copy old tutorials that call client.search() Many articles use client.search(), client.recommend() or client.discover(). Those wrap the older endpoints /points/search, /points/recommend and /points/discover, which Qdrant deprecated in version 1.13.3 and dropped from its API documentation in 1.19. Use query_points and the Query API shown here. In the same way, the helper upload_records is deprecated; use upload_points or upload_collection for bulk loading.
Try it
  1. Save the script, run it with python cities.py, and compare the output with your curl results.
  2. Run it a second time and read the error. Explain why the collection already exists.
  3. Remove the filter from the query and see how the ranking changes.

Everyday operations, grouped by what you are trying to do

All requests below are REST calls against the cities collection. Each has an equivalent in the clients and in gRPC.

Manage collections

Task Call
List all collections GET /collections
Does it exist? GET /collections/cities/exists
Inspect configuration and status GET /collections/cities
Delete it (and every point in it) DELETE /collections/cities

GET /collections/cities is the one to learn. Its response shows status, points_count, indexed_vectors_count, and the configuration. The status is a traffic light: green means ready, yellow means the optimizer is busy rebuilding something (searches still work), grey means optimizations are pending but paused, which usually follows a restart, and red means an unrecoverable error.

Collection names are plain strings. Since 1.19.1, the names . and .. are rejected, and a vector size above 65,536 or an empty vector is rejected too.

Fetch, list, and count points

To fetch specific points by ID:

BASH
curl -X POST 'http://localhost:6333/collections/cities/points' \
  -H 'Content-Type: application/json' \
  -d '{"ids": [1, 2], "with_payload": true, "with_vector": false}'

To page through all points without any similarity, use scroll:

BASH
curl -X POST 'http://localhost:6333/collections/cities/points/scroll' \
  -H 'Content-Type: application/json' \
  -d '{"limit": 2, "with_payload": true}'

The response includes a next_page_offset. Pass it back as "offset" in the next request to get the next page. When next_page_offset is null, you have reached the last page. Scroll accepts a filter too, so it is the right tool for "give me every point where country is Egypt", an exact listing with no ranking.

To count points:

BASH
curl -X POST 'http://localhost:6333/collections/cities/points/count' \
  -H 'Content-Type: application/json' \
  -d '{"filter": {"must": [{"key": "country", "match": {"value": "Egypt"}}]}, "exact": true}'

Counts are approximate unless you pass "exact": true. The same note applies to the points_count and indexed_vectors_count fields in collection info: they are good for trends, not for accounting.

Update and delete points

Upserting an existing ID replaces the whole point: the vector and payload you send become the new truth. If you send a point with only a payload change and leave out the vector, you have not edited the payload; you have overwritten the record. Updating individual payload keys without touching vectors uses the dedicated payload endpoints, which you meet at the mid level. For now, the safe beginner rule is: when you upsert, send the complete point.

Deleting points by ID or by filter:

BASH
curl -X POST 'http://localhost:6333/collections/cities/points/delete?wait=true' \
  -H 'Content-Type: application/json' \
  -d '{"points": [6]}'
BASH
curl -X POST 'http://localhost:6333/collections/cities/points/delete?wait=true' \
  -H 'Content-Type: application/json' \
  -d '{"filter": {"must": [{"key": "budget", "match": {"value": "high"}}]}}'

Deletes are "soft" internally: Qdrant marks points as removed and the optimizer reclaims the space later. From your side the points vanish immediately from search.

Search by an existing point

Sometimes you want "more like this stored item" without having its vector in hand. Pass the point ID as the query:

BASH
curl -X POST 'http://localhost:6333/collections/cities/points/query' \
  -H 'Content-Type: application/json' \
  -d '{"query": 2, "limit": 3, "with_payload": true}'

This asks for the three points nearest to point 2 (Alexandria). The stored point itself is excluded from its own results.

Try it
  1. Scroll the cities collection two points at a time and follow next_page_offset until it is null.
  2. Count the mid-budget cities with "exact": true.
  3. Delete one city by ID, count again, then upsert it back.

Filtering in depth

Filters are what separate a vector database from a vector library, and Qdrant's filter language is expressive. A filter is a JSON object built from three clauses and many conditions.

The clauses combine conditions:

  • must means every condition has to match (logical AND).
  • should means at least one has to match (logical OR).
  • must_not means none of them may match.

You can use more than one clause at a time and nest filters inside each other. This request finds mid- or low-budget cities that are not in Egypt:

JSON
{
  "filter": {
    "must": [{"key": "budget", "match": {"any": ["low", "mid"]}}],
    "must_not": [{"key": "country", "match": {"value": "Egypt"}}]
  }
}

The most common conditions are:

  • match with value for an exact value, any for "one of these values", and except for "none of these".
  • range with gt, gte, lt, lte for numbers, and for dates written as strings.
  • match with text for full-text matching on a field that has a text index. (Qdrant also has text_any, phrase, and, since 1.19, prefix matching, which the mid-level guide covers.)
  • geo_radius, geo_bounding_box, and geo_polygon for location.
  • is_empty, is_null, values_count, has_id, and has_vector for structure checks.
  • nested for filtering inside arrays of objects.

Nested payload keys use dot notation, for example "key": "address.city", and arrays of objects use a[].b.

Here is the Python equivalent of a two-condition filter, to show how the objects line up with the JSON:

PYTHON
from qdrant_client.models import Filter, FieldCondition, MatchValue, MatchAny

flt = Filter(
    must=[FieldCondition(key="budget", match=MatchAny(any=["low", "mid"]))],
    must_not=[FieldCondition(key="country", match=MatchValue(value="Egypt"))],
)

Filters apply identically to search, scroll, count, and delete. You learn the language once and reuse it everywhere. That consistency is one of the nicest things about the API.

There are two behaviours to be aware of. First, a filter on a field that does not exist in a point's payload simply does not match that point; there is no error. A typo in "key": "contry" quietly returns nothing, which makes this a classic cause of "my filter returns empty results". Second, the type matters: a number stored as the string "5" does not match the integer 5. Decide on types when you design the payload and keep them consistent.

Filter in Qdrant

  • Applied while searching, so the top results are chosen from allowed points only
  • You ask for 3 results and you get 3 (if 3 allowed points exist)

Filter after the search, in your own code

  • You fetch the top 3 and then throw some away
  • You may be left with 0 or 1 result even though many allowed points exist
Try it
  1. Write a filter for "budget is mid, and country is not Jordan" and run it with a search.
  2. Deliberately misspell a payload key and confirm the result is empty rather than an error.
  3. Add a range filter. First add a numeric payload field such as rating to two points by upserting them again with complete data.

Payload indexes: the habit that pays off first

By default a payload field is not indexed. A filter on it still works, but Qdrant has to inspect points' payloads one by one, which is fine at six points and painful at six million. A payload index is a per-field index that makes filtering on that field fast. It also helps the query planner, the part of Qdrant that decides whether to walk the HNSW graph or scan a small filtered subset directly.

Create one with a single call. The field type is one of keyword (exact strings), integer, float, bool, geo, datetime, text (full-text), or uuid:

BASH
curl -X PUT 'http://localhost:6333/collections/cities/index?wait=true' \
  -H 'Content-Type: application/json' \
  -d '{"field_name": "budget", "field_schema": "keyword"}'

There is a subtlety that costs beginners real time. Qdrant builds extra graph connections for each indexed payload field, so that filtered searches do not get cut off from parts of the graph. Those connections are only created for payload indexes that exist when the graph is built. So the rule is:

Create payload indexes before you load your data, for every field you will filter on.

If you add an index after the data is in, it still speeds up filtering, but the graph does not get the extra connections until it is rebuilt. Creating the indexes first, while the collection is still empty, avoids the problem entirely. Do not index every field "just in case"; each index costs memory and slows writes. Index the fields that appear in your filters, and leave the rest.

Another reason to index: Qdrant has an optional strict mode that can reject requests which filter on unindexed fields instead of running them slowly. It is turned on by default on Qdrant Cloud. A request rejected for filtering on an unindexed field is a signal that you forgot an index, not a bug.

Decide your filter fields when you design the collection Before writing ingestion code, list the fields your users will filter on (tenant, category, language, date). Create a payload index for each, then upload. That ordering is the cheapest performance fix you will ever make.
Try it
  1. Create a keyword index on country in the cities collection.
  2. Run GET /collections/cities and find the payload schema section that now lists the field.
  3. Write down which fields you would index for a support-ticket collection.

Named vectors and sparse vectors, briefly

You will meet two variations soon, so a short introduction avoids surprise.

Named vectors let one point hold several vectors, each with its own size and metric. Imagine a product with an image vector from a vision model and a text vector from a language model. You declare them at creation:

JSON
{"vectors": {
  "image": {"size": 4, "distance": "Dot"},
  "text":  {"size": 8, "distance": "Cosine"}
}}

When searching, say which vector space to use with "using": "image". A collection created with a single unnamed vector (as in our project) has just the default unnamed space, and you do not need using.

Sparse vectors store only the positions and values of non-zero entries, as {"indices": [1, 42], "values": [0.22, 0.8]}. They represent keyword-style signals such as BM25 or SPLADE weights, where a vocabulary of tens of thousands of terms has only a few non-zero per document. A sparse vector must be named and always uses the Dot metric. You declare the space as {"sparse_vectors": {"text": {}}}, adding "modifier": "idf" for BM25-style weighting.

Putting a dense and a sparse vector on the same point and searching both is called hybrid search, and it often beats either alone because dense vectors capture meaning while sparse ones catch exact terms and rare words. It uses the Query API's prefetch feature and is a mid-level topic. What matters today is that you know the option exists and that your collection design leaves room for it: if you suspect you will want hybrid search, name your vectors from the start.

You can add vectors later Since version 1.18 you can add or remove a named vector on an existing collection with PUT and DELETE on /collections/{name}/vectors/{vector_name}. You still cannot change the length of an existing space, so a new embedding model means a new vector name or a new collection.
Try it
  1. Create a collection with two named vectors using the JSON above and insert one point with both vectors.
  2. Search it twice, once with "using": "image" and once with "using": "text", using query vectors of the right lengths.
  3. Try a query vector of the wrong length and read the error message.

Configuration, and how Qdrant starts

Out of the box Qdrant needs no configuration. When you do need it, there are two routes and one rule about which wins.

Qdrant reads settings from several layers, lowest priority first: built-in defaults, then config/config.yaml, then a file named after the run mode (the Docker image uses production, so config/production.yaml), then config/local.yaml, then a file passed with --config-path, and finally environment variables, which always win. YAML, TOML, JSON, and INI files are all accepted.

Environment variables use the prefix QDRANT__ and a double underscore for each level of nesting. These are the settings a beginner is most likely to touch:

Setting Environment variable Default What it does
service.http_port QDRANT__SERVICE__HTTP_PORT 6333 REST and Web UI port
service.grpc_port QDRANT__SERVICE__GRPC_PORT 6334 gRPC port (null disables)
service.host QDRANT__SERVICE__HOST 0.0.0.0 Which network interface to listen on
service.api_key QDRANT__SERVICE__API_KEY none Require this key on requests
log_level QDRANT__LOG_LEVEL INFO How chatty the logs are
telemetry_disabled QDRANT__TELEMETRY_DISABLED false Whether anonymous usage statistics are sent
storage.storage_path QDRANT__STORAGE__STORAGE_PATH ./storage Where data lives (in Docker: /qdrant/storage)

For example, to run on a different log level and with telemetry off:

BASH
docker run -p 6333:6333 -p 6334:6334 \
  -e QDRANT__LOG_LEVEL=DEBUG \
  -e QDRANT__TELEMETRY_DISABLED=true \
  -v "$(pwd)/qdrant_storage:/qdrant/storage:z" \
  qdrant/qdrant:v1.19.1

If you prefer a file, mount one at /qdrant/config/production.yaml:

production.yaml
log_level: INFO
service:
  http_port: 6333
telemetry_disabled: true

The configuration is validated at start-up, and an invalid value aborts the launch with a message that names the key, such as invalid type: 64-bit integer -1, expected an unsigned 64-bit or smaller integer for key storage.hnsw_index.max_indexing_threads in config/production.yaml. That is helpful: read the key name in the message, fix the value, and start again.

If you use Docker Compose, a minimal file looks like this:

docker-compose.yml
services:
  qdrant:
    image: qdrant/qdrant:v1.19.1
    restart: always
    container_name: qdrant
    ports:
      - "127.0.0.1:6333:6333"
      - "127.0.0.1:6334:6334"
    volumes:
      - ./qdrant_data:/qdrant/storage

Notice the 127.0.0.1: in front of the ports. It makes Qdrant reachable only from your own machine, which is the safe default for development. The next section explains why.

Many settings that tune collections (HNSW parameters, optimizer thresholds, quantization) are set per collection when you create or later update it, not in the config file. Qdrant Cloud does not allow editing the server config at all, since it manages that for you.

Try it
  1. Restart Qdrant with QDRANT__LOG_LEVEL=DEBUG and compare the logs with before.
  2. Set an invalid value, such as a negative number for a port, and read the start-up error.
  3. Write the Compose file above and start it with docker compose up.

Securing a fresh server

This section is short but important. A self-hosted Qdrant starts with no authentication, no encryption, and listening on every network interface. The documentation says plainly that this default is not production-ready. If you run it on a cloud virtual machine with port 6333 open to the internet, anyone can read, change, or delete your collections. People have lost data this way.

Three habits keep you safe from day one.

Bind to localhost while developing. Publish ports as -p 127.0.0.1:6333:6333, as in the Compose file above. Only processes on your own machine can connect.

Set an API key before exposing the server to any network. Set the environment variable QDRANT__SERVICE__API_KEY to a long random secret. Clients then send it in a header named api-key:

BASH
curl -H 'api-key: YOUR_SECRET' http://localhost:6333/collections

In Python you pass api_key="YOUR_SECRET" to QdrantClient. Health endpoints (/healthz, /livez, /readyz) remain open so probes keep working. Keep the key out of your Git repository: inject it from your deployment platform's secret store, not from a committed file. There is also a read_only_api_key for clients that should only read.

Never expose port 6335. That is Qdrant's internal peer-to-peer port, used only between cluster nodes, and it is not protected by the API key. Do not publish it at all unless you are running a cluster, and then keep it on a private network.

API keys protect against casual access but travel in plain text unless you also enable TLS, which a beginner usually delegates to a reverse proxy or to Qdrant Cloud. Mid-level and Senior cover TLS, JWT tokens with per-collection permissions, and network policies. For a student project, "localhost only, or an API key plus a firewall" is the right amount.

An open port is a public database Cloud virtual machines on public IPs are scanned constantly. Treat "Qdrant with no API key reachable from the internet" as an incident waiting to happen. If you are experimenting, bind to 127.0.0.1.
Try it
  1. Restart Qdrant with -e QDRANT__SERVICE__API_KEY=test-key-123.
  2. Call /collections without a key and read the rejection, then call it again with the api-key header.
  3. Call /healthz without a key and confirm it still answers.

Common errors and how to read them

Almost every Qdrant error comes from one of a short list of causes. Learn to read the message before searching the web.

Wrong vector length. Upserting or searching with a vector whose length does not match the collection returns an error that names the expected dimension. The collection says 768, you sent 384. The cause is nearly always a different embedding model than the one you used to create the collection. Fix the model or create a new collection.

Collection not found or already exists. Not found: Collection ... doesn't exist means a typo in the name, or you are talking to a different server than you think. The reverse, creating a collection that exists, is refused rather than overwritten. Check with GET /collections/{name}/exists.

Search returns nothing even though data exists. Work through the list. Did the write finish (use wait=true)? Is your filter key spelled correctly and of the right type? Is the collection actually empty (points_count)? Are you searching a named vector without saying "using"? Did you apply score_threshold on a Euclid collection with a similarity-style number?

Too many files open (OS error 24). Each segment holds open files, and a large collection can exceed the default operating-system limit. Raise it: docker run --ulimit nofile=10000:10000 qdrant/qdrant, or ulimit -n 10000 before starting the binary.

Filesystem warnings and panics. A message like Filesystem check failed ... FUSE filesystems may cause data corruption, or a panic mentioning OutputTooSmall, means the data folder is on a filesystem Qdrant cannot trust, such as a Windows bind mount under WSL, FUSE, or NFS. Use a named Docker volume or a local disk.

Can't open Collections meta Wal ... WouldBlock. Two Qdrant processes share one storage folder. Stop one, or give each its own folder.

Grey collection status. The optimizer has pending work but is paused, typically after a restart in the middle of optimization. Send any update to wake it, for instance PATCH /collections/{name} with {"optimizers_config": {}}, or press "Trigger Optimizers" in the Web UI.

indexed_vectors_count is 0 or lower than points_count. Not an error. Small segments are not indexed, as described earlier. Counts in collection info are approximate; use the count API with "exact": true for real numbers.

HTTP 507 or ResourceExhausted on writes. Since 1.19 Qdrant has resource quotas. When disk usage passes the configured limit it refuses writes with a message such as Disk usage is at 95% of total capacity, exceeding the configured limit of 90%. Reads keep working. Deleting points, adding disk, or raising the limit fixes it.

HTTP 429. A rate limit from strict mode was exceeded. Slow down and retry after the delay in the response.

Duplicates or gaps when paging search results. Using offset with an approximate search can show the same point twice or skip one. This is expected with approximate search. For complete, ordered listings use scroll; for search pages, fetch more results and page in your own code.

A good diagnostic routine, in order: read the exact message, check GET /collections/{name} for status and counts, then read the server log.

Try it
  1. Cause three errors on purpose: a wrong vector length, a missing collection name, and a duplicate collection creation. Read each message.
  2. Write one sentence for each error saying what it means and how you would fix it.

Putting it all together: a small search service

Here is a complete small project that follows the habits from this guide: pinned version, localhost binding, collection design before data, payload index before upload, batched upsert, and a filtered search. It uses invented four-number vectors so you can run it without an embedding model. In real work, the only line that changes is where the vectors come from.

Start the server bound to localhost:

BASH
docker run -d --name qdrant \
  -p 127.0.0.1:6333:6333 -p 127.0.0.1:6334:6334 \
  -v "$(pwd)/qdrant_storage:/qdrant/storage:z" \
  qdrant/qdrant:v1.19.1

Then the script:

tour_guide.py
from qdrant_client import QdrantClient
from qdrant_client.models import (
    Distance, VectorParams, PointStruct, PayloadSchemaType,
    Filter, FieldCondition, MatchValue,
)

NAME = "tour_guide"
client = QdrantClient(url="http://localhost:6333")

# 1. Start clean so the script can be re-run.
if client.collection_exists(NAME):
    client.delete_collection(NAME)

# 2. Design: vector length and metric are decided up front.
client.create_collection(
    collection_name=NAME,
    vectors_config=VectorParams(size=4, distance=Distance.COSINE),
)

# 3. Index the fields we will filter on BEFORE loading data.
client.create_payload_index(NAME, field_name="budget", field_schema=PayloadSchemaType.KEYWORD)
client.create_payload_index(NAME, field_name="country", field_schema=PayloadSchemaType.KEYWORD)

# 4. Load the data (a placeholder for your embedding step).
cities = [
    (1, [0.1, 0.9, 0.7, 0.1], "Cairo", "Egypt", "low"),
    (2, [0.9, 0.7, 0.4, 0.4], "Alexandria", "Egypt", "low"),
    (3, [0.7, 0.2, 0.9, 0.3], "Dubai", "UAE", "high"),
    (4, [0.8, 0.6, 0.3, 0.8], "Muscat", "Oman", "mid"),
    (5, [0.8, 0.5, 0.6, 0.4], "Casablanca", "Morocco", "mid"),
    (6, [0.0, 0.8, 0.5, 0.6], "Amman", "Jordan", "mid"),
]
client.upsert(
    collection_name=NAME,
    wait=True,
    points=[
        PointStruct(id=i, vector=v, payload={"city": c, "country": k, "budget": b})
        for i, v, c, k, b in cities
    ],
)

# 5. Search with a filter.
def recommend(preference, budget, limit=3):
    result = client.query_points(
        collection_name=NAME,
        query=preference,
        query_filter=Filter(must=[FieldCondition(key="budget", match=MatchValue(value=budget))]),
        limit=limit,
        with_payload=True,
    )
    return [(p.payload["city"], round(p.score, 3)) for p in result.points]

print(recommend([0.9, 0.6, 0.3, 0.5], "mid"))
print(recommend([0.1, 0.9, 0.6, 0.2], "low"))

Walk through what each stage teaches. Stage 1 makes the script repeatable, which matters because creation is refused when a collection already exists. Stage 2 fixes size and metric. Stage 3 creates payload indexes while the collection is empty, so the graph will get its filter-aware connections. Stage 4 uploads in one batch; for large loads you would batch 64 to 256 points per request, and use the client's upload_points helper with parallelism. Stage 5 wraps a filtered query in a function, which is the shape a web endpoint would have.

To turn this into a real application, replace the made-up vectors with an embedding model's output. Use the same model, with the same length, for the stored documents and for each query; put the original text in the payload so you can show it, because Qdrant does not keep your input text unless you store it yourself. Wrap recommend in a small FastAPI endpoint, and if the application is a question-answering system over documents, look at how LangChain and LlamaIndex use Qdrant as a vector store, and at RAGAS for measuring whether the retrieval is good.

Finally, take the project one step toward a real deployment: stop the container (docker stop qdrant), start it again (docker start qdrant), and run only the search part. The collection is still there because the storage folder survived the container. That is persistence working as designed.

Try it
  1. Run the script and check the output matches your reasoning about the toy data.
  2. Stop and restart the container, then confirm the data survived by calling GET /collections/tour_guide.
  3. Extend the toy dataset with four more cities and a rating payload, add a range filter, and add an index for rating as integer before you load.

What you can now do, and what comes next

You can now explain what a vector database is for and where Qdrant sits between an embedding model and your application. You know the four nouns (collection, point, vector, payload) and the first-day vocabulary around them (segment, WAL, optimizer, shard, alias). You can pick a distance metric, run Qdrant in Docker with persistent storage, verify it with a health check, and read its version. You created a collection, upserted points, searched with the Query API, filtered on payload, scrolled, counted, and deleted. You know why payload indexes come before ingestion, how to configure Qdrant through environment variables, how to lock a fresh server down, and how to read the common errors.

The habits worth keeping are few. Pin the version. Decide vector length and metric from the embedding model before creating anything. Create payload indexes before uploading data. Add wait=true to curl writes. Bind to localhost or set an API key. Use the Query API rather than the deprecated search endpoints. Store the original text or a link back in the payload.

What comes next is in the mid-level guide: updating payloads and vectors precisely, batch operations, the update_mode and conditional-update options, facets and grouping, hybrid search with dense and sparse vectors, fusion, recommendation and discovery queries, full-text indexes, quantization to shrink memory, memory tiers, snapshots and aliases for safe re-indexing, and monitoring. The senior guide covers clusters (shards, replicas, consensus), capacity planning, multi-tenancy, security in depth, upgrades (never skip a minor version), and when to choose something other than Qdrant. If you want to compare approaches, read the guides on Pinecone, Weaviate, Milvus, Chroma, and pgvector, and then decide on your own data which fits.

Sources