Skip to content
Back to student guides
MilvusLLMsVector databases3 levels98 sectionsCovers Milvus 3.0

The Complete Milvus Guide

Run billion-scale vector search with the cloud-native Milvus database. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

Official docs AI-drafted · community review in progressHelp review it
15sections
34examples

This is part one of three. It covers everything you need to do real work with Milvus as a beginner, not a teaser. By the end you can start a Milvus server on your laptop (or skip the server entirely and run it inside a Python script), design a collection, store embeddings with their metadata, search them by meaning, filter by scalar fields, update and delete data, read the common errors, and build a small semantic search project from scratch. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. Vector search feels abstract until you have watched your own query return the right sentence, and these ideas only stick once you have also watched a search fail because the collection was not loaded.

This guide targets Milvus 3.0.2, the current stable release at the time of writing, with the matching Python client, pymilvus 3.0.2. A great deal of what you find online was written for Milvus 2.5 or earlier. We flag the places where that older material is now wrong.

What Milvus is, and the problem it solves

Milvus is an open-source vector database. It stores lists of numbers called vectors (also called embeddings), together with ordinary fields such as a title or a date, and it answers one question very quickly: which stored vectors are closest to this one?

To see why that matters, start with what an embedding is. A machine learning model can turn a piece of text, an image, or a sound clip into a list of a few hundred or a few thousand numbers. The model is trained so that things with similar meaning end up with similar numbers. The sentence "How do I reset my password?" and the sentence "I forgot my login credentials" share almost no words, yet an embedding model places them close together. "The quarterly revenue grew" lands far away from both. Closeness in number space stands in for closeness in meaning.

That gives you a new kind of search. A traditional database finds rows where a column equals a value, or where a text field contains a word. A vector database finds rows whose meaning is near the meaning of your query, even when no word matches. This is the engine behind semantic search, recommendation, duplicate detection, image search, and the retrieval step of retrieval-augmented generation (RAG), where a language model is given relevant passages pulled from your own documents before it answers.

YOUR DATAtext, images, audio
→
EMBEDDING MODELdata to numbers
→
MILVUSstore and index vectors
→
NEAREST NEIGHBOURSthe most similar items

Note what Milvus does and does not do. Milvus does not create embeddings for you in the simple case. An embedding model does that (a model you run, or a hosted service), and you hand the resulting vectors to Milvus. Milvus stores them, builds an index over them, and searches them. Newer versions can also call an embedding provider from the server side, which is a mid-level topic. For now, the mental split is: the model understands meaning, Milvus finds neighbours.

What came before

You could always do this with brute force: keep every vector in a NumPy array, compute the distance from your query to each one, sort, and take the top few. For ten thousand vectors that is fine and you should not feel bad about it. At ten million vectors it is too slow, and it also leaves you to build everything around it: persistence, metadata filtering, updates and deletes, access from several services, backups.

Libraries such as FAISS solved the speed part by providing fast approximate-nearest-neighbour indexes, but a library is not a database. It lives inside your process and leaves storage, filtering, and concurrent access to you. Milvus is the database layer around that idea: a server (or an embedded library) that persists data, indexes it, filters it, scales, and speaks a network protocol that many programs can use. There is a dedicated guide to FAISS if you want to see the library underneath that world.

The three ways to run Milvus

Milvus comes in three deployment modes, and picking the right one is your first decision.

  • Milvus Lite is a Python library. You install it with pip, point it at a file such as ./milvus_demo.db, and it runs inside your Python process. There is no server and no Docker. It is for learning and prototyping, and it runs on Linux and macOS but not on Windows.
  • Standalone is a single Milvus process in one container, plus two companions it depends on (etcd for metadata and an object store for data files). It is the right choice for a laptop, a small server, or a first real project.
  • Distributed runs Milvus as several cooperating services on Kubernetes. It is for large data and high traffic, and it is covered in the Mid-level and Senior parts.

There is also Zilliz Cloud, the managed service run by the company behind Milvus, which gives you the same API without operating anything yourself. The code you write in this guide works against all of these with a one-line change to the connection address, which is a real advantage when you start small and grow.

Try it
  1. Write three sentences about different topics, for example a cooking tip, a football result, and a database tip.
  2. Write a fourth sentence that is related to one of them but shares no important words with it.
  3. Decide by hand which of the first three the fourth is closest to.
you can do this easily, because meaning is something you perceive. An embedding model's job is to produce numbers that preserve exactly that judgement, and Milvus's job is to find the nearest neighbours among millions of such numbers.

The core ideas: five nouns

Milvus has a larger vocabulary than most beginner tools, but five nouns carry nearly all of it. Learn these and the documentation becomes readable.

DATABASEa namespace
→
COLLECTIONlike a table
→
FIELDlike a column
→
ENTITYlike a row
→
INDEXa search shortcut

Collection. A collection is the unit you work with, and it is the closest thing to a table. You create one, insert into it, search it, and drop it. A collection has a schema that says what fields every record contains. A database is just a namespace that groups collections. There is always one called default, and a beginner can work in it forever.

Field and schema. A schema lists the fields. Every collection needs exactly one primary key field (an integer or a string) that uniquely identifies each record, and at least one vector field that holds the embedding. It can also have any number of scalar fields, which is Milvus's word for ordinary values: text, numbers, booleans, JSON, arrays. The vector field has a dimension, the length of the list, and every vector you store in that field must have exactly that length. If your embedding model produces 768 numbers, the field has dimension 768, and a vector of 767 numbers is rejected.

Entity. An entity is one record: a row. In Python you write it as a dictionary, for example {"id": 1, "vector": [...], "text": "..."}.

Index. An index is a data structure built over a vector field so that searches do not have to compare against everything. Without an index Milvus can only do an exact comparison with every vector (a brute-force scan). With one it does an approximate search: dramatically faster, with a small, tunable chance of missing a true neighbour. Approximate nearest-neighbour search is the central trade-off of the whole field, and we return to it below.

Metric type. To find "closest", Milvus needs a definition of distance. For ordinary float vectors the three metrics are COSINE (the angle between vectors, ignoring their length), L2 (straight-line Euclidean distance), and IP (inner product). The rule is to use the metric your embedding model was trained for. For most text embedding models that is cosine similarity, and cosine is also the default for the quick setup you will use in a moment.

Higher or lower is better? It depends on the metric With L2 a smaller number means closer, so the best hit has the smallest distance. With COSINE and IP a larger number means more similar, so the best hit has the largest. Milvus handles the ordering for you and returns the best results first, but the field it returns is called distance in every case. If you print a score and it looks "backwards", check your metric before assuming a bug.

Load and release, the idea that surprises everyone

Here is the concept that trips up almost every newcomer. For a search to work, the collection must be loaded. Loading means the data and its index are brought into memory on the server so that queries can be answered fast. A collection that exists but is not loaded can be inserted into, but cannot be searched, and you get the error collection not loaded.

The reassuring part: when you create a collection with the quick setup (shown below), or when you pass index parameters at creation time, Milvus loads it for you. You meet load_collection only when you take a different route, or after you release a collection to free memory. Remember the pairing, though, because it is the first thing to check when a search fails.

Segments, in one paragraph

You do not need to manage them, but you will see the word in logs. Milvus stores incoming data in segments. New writes land in a growing segment. When a segment is big enough or you ask for a flush, it is sealed, turned into immutable files in object storage, and indexed. Searches look across both. This is why a freshly inserted row can occasionally be invisible for a moment and why the consistency level, discussed later, exists.

Try it
  1. Imagine a collection of support articles. Write down its primary key, its vector field (with a dimension of 384), and three scalar fields.
  2. For each scalar field, decide whether you would ever filter by it.
  3. Say which metric you would choose if your model's documentation says it was trained with cosine similarity.
a schema such as id (integer key), vector (dimension 384), title (text), category (text), year (integer), with the metric set to COSINE. Fields you filter on belong in the schema; they are what makes a search "find similar articles, but only from 2025".

Approximate nearest-neighbour search, in plain words

Before touching code, spend two minutes on the idea that explains almost every setting you will meet later, because it also explains the interview questions.

Searching a million vectors exactly means a million distance calculations per query. Approximate search avoids most of them by organising the vectors in advance. Two families dominate. Graph indexes (the best known is HNSW) link each vector to some of its near neighbours and then walk that graph from a starting point toward the query, always stepping to a closer vertex. Cluster indexes (the IVF family) split the space into regions around centre points, then search only the few regions nearest the query. Both examine a small fraction of the data and return answers that are almost always the true nearest ones.

The word "almost" has a name: recall, the fraction of true nearest neighbours that the search actually returns. Higher recall costs more time and memory. Each index has parameters that move you along that line, and the defaults are sensible. As a beginner you will not tune them. You will let Milvus choose with the setting called AUTOINDEX, which picks an index and parameters for you. That is what the quick setup does.

Keep two facts. First, an index is built over a vector field and you must have one before searching at scale. Second, FLAT is the one index that is exact: it compares against everything and never misses. It is the right choice for small datasets and for checking what the approximate indexes are missing. Milvus Lite, in fact, always uses FLAT regardless of what you ask for, which is fine because it is meant for small data.

Installing and checking the setup

You have three options. Start with Milvus Lite if you only want to learn the API, because it needs nothing but Python. Move to Docker standalone when you want the real server, the visual tools, and a setup that resembles production.

Option one: Milvus Lite (Linux and macOS)

BASH
pip install -U "pymilvus[milvus-lite]"

That single command installs the Python client and the embedded engine. It supports Ubuntu 20.04 and later (x86_64 and arm64) and macOS 11 and later (Apple Silicon and Intel). It does not run on Windows. If you try, you get ModuleNotFoundError: No module named 'milvus_lite'. Windows users should use WSL 2 (a Linux environment inside Windows) or Docker, both described below.

Connecting is one line, and the "address" is just a file path:

PYTHON
from pymilvus import MilvusClient

client = MilvusClient("./milvus_demo.db")
print(client.list_collections())
TEXT
[]

An empty list means the connection works and nothing is stored yet. The file milvus_demo.db appears next to your script and holds your data between runs. Milvus Lite has real limits: it does not support partitions, users and roles, or aliases, and it ignores some collection options. It is a learning tool, and a pleasant one. Later, milvus-lite dump can export a collection from the file so you can load it into a full Milvus.

Option two: Milvus standalone in Docker

Standalone is the real server. You need a machine with at least 8 GB of memory available (16 GB is recommended) and a CPU with SIMD instructions (SSE4.2, AVX, AVX2, or AVX-512), which any recent laptop has. On Linux you can check with:

BASH
lscpu | grep -e sse4_2 -e avx -e avx2 -e avx512

On a Mac, Docker Desktop's virtual machine should be given at least 2 virtual CPUs and 8 GB of memory, which you set in its settings. On Windows, use Docker Desktop with WSL 2 and keep your data in the Linux filesystem rather than the Windows one. If you have never used Docker, the Docker guide in this series covers containers from zero, and this guide assumes only that docker runs on your machine.

The quickest start is the official script, which launches a single container with an embedded etcd and local storage:

BASH
curl -sfL https://raw.githubusercontent.com/milvus-io/milvus/master/scripts/standalone_embed.sh -o standalone_embed.sh
bash standalone_embed.sh start

The script creates a container named milvus-standalone. Milvus listens for client connections on port 19530, and a small web interface and the metrics endpoint listen on port 9091. Data is stored in a volumes/milvus folder beside the script. The same script accepts stop, restart, delete, and upgrade. On Windows there is an equivalent standalone_embed.bat, run from PowerShell with Docker Desktop started as administrator.

The other official route is Docker Compose, which runs three containers: Milvus, plus separate etcd and MinIO (an S3-compatible object store) containers. It is closer to production and shows you the moving parts:

BASH
wget https://github.com/milvus-io/milvus/releases/download/v3.0.2/milvus-standalone-docker-compose.yml -O docker-compose.yml
sudo docker compose up -d
sudo docker compose ps

docker compose ps should list milvus-standalone, milvus-minio, and milvus-etcd as running (healthy). The file name differs slightly between documentation pages (.yaml on one, .yml on another), so if the download fails, check the release assets for the exact name. To tear everything down and delete the data, run sudo docker compose down and then sudo rm -rf volumes. Be careful: that deletes your stored vectors.

Check the web interface before writing any code Open http://127.0.0.1:9091/webui/ in a browser. If a page loads, the server is up and you have ruled out a whole category of connection problems before your Python script ever runs. The same port serves Prometheus metrics at /metrics, which you will care about later.

Installing the Python client

BASH
pip install -U pymilvus
python -c "import pymilvus; print(pymilvus.__version__)"

For Milvus 3.0.2 you want pymilvus 3.0.2. Use a virtual environment so the client does not collide with other projects. There is also an optional extra, pip install "pymilvus[model]", that adds helper classes for generating embeddings. We will not need it, but you will see it in tutorials.

Now connect to your standalone server and prove the whole chain works:

PYTHON
from pymilvus import MilvusClient

client = MilvusClient(uri="http://localhost:19530", token="root:Milvus")
print(client.list_collections())
TEXT
[]

The token here is user:password. A fresh Milvus installation has authentication turned off, so the token is ignored, but the documented default account is user root with password Milvus, and including it makes your code ready for the day authentication is on. We return to the security consequences in a later section.

Two version traps in tutorials Older tutorials show from pymilvus import connections, Collection and calls such as connections.connect() and Collection(...). That older object-oriented style still exists in places, but the current documentation is built around MilvusClient, and so is this guide. Also, any architecture diagram that shows separate RootCoord, QueryCoord, DataCoord, or IndexNode services is from before Milvus 2.6, which merged and removed them. Do not copy commands that scale an "index node".
Try it
  1. Pick Milvus Lite (if you are on Linux or macOS and just want the API) or Docker standalone.
  2. Install pymilvus and connect with MilvusClient.
  3. Call client.list_collections().
  4. If you used Docker, also open http://127.0.0.1:9091/webui/.
an empty list [] and, for Docker, a loading web page. If the connection hangs or is refused, the container is not running yet (check docker ps) or is still starting; give it a minute and look at docker logs milvus-standalone.

Your first collection in four calls

The fastest path from nothing to a working vector search is the quick setup. In this section we use tiny three-number vectors so you can read every value; the next sections use realistic ones.

Create the collection

PYTHON
from pymilvus import MilvusClient

client = MilvusClient("./milvus_demo.db")   # or uri="http://localhost:19530"

if client.has_collection("demo_collection"):
    client.drop_collection("demo_collection")

client.create_collection(
    collection_name="demo_collection",
    dimension=4,
)

Two arguments are all it takes. Behind that short call, Milvus does a surprising amount, and you should know what, because you did not ask for it:

  • It creates a primary key field named id (an integer) and a vector field named vector with dimension 4.
  • It turns on the dynamic field, which means you can insert extra keys that are not in the schema, and they are stored in a hidden JSON field.
  • It builds an index on the vector field using AUTOINDEX, with the default metric COSINE.
  • It loads the collection, so it is searchable immediately.

The quick setup is excellent for learning and for simple projects. Its cost is flexibility: field names are fixed, and the primary key is not auto-generated unless you pass auto_id=True. The custom schema in a later section removes those limits.

Insert some data

PYTHON
data = [
    {"id": 0, "vector": [0.9, 0.1, 0.0, 0.1], "text": "Cats purr when they are content.", "topic": "animals"},
    {"id": 1, "vector": [0.8, 0.2, 0.1, 0.0], "text": "Dogs need daily walks.", "topic": "animals"},
    {"id": 2, "vector": [0.1, 0.9, 0.2, 0.0], "text": "Python is a programming language.", "topic": "tech"},
    {"id": 3, "vector": [0.0, 0.8, 0.3, 0.1], "text": "Milvus stores vectors.", "topic": "tech"},
]

result = client.insert(collection_name="demo_collection", data=data)
print(result)
TEXT
{'insert_count': 4, 'ids': [0, 1, 2, 3], 'cost': 0}

Each entity is a dictionary. The id and vector keys match the schema. The text and topic keys are not in the schema at all; they were accepted because the dynamic field is on, and you can search and filter on them just like declared fields. The return value tells you how many rows were inserted and which ids they received.

PYTHON
query_vector = [0.85, 0.15, 0.05, 0.05]   # a "cat-like" direction

results = client.search(
    collection_name="demo_collection",
    data=[query_vector],
    limit=2,
    output_fields=["text", "topic"],
)

for hit in results[0]:
    print(hit["id"], round(hit["distance"], 3), hit["entity"])
TEXT
0 0.999 {'text': 'Cats purr when they are content.', 'topic': 'animals'}
1 0.995 {'text': 'Dogs need daily walks.', 'topic': 'animals'}

Your exact decimals may differ slightly, but the shape is what matters. Read it closely, because it is the most common thing beginners misread:

  • data is a list of query vectors, even when you have one. You can search several vectors in one call.
  • The result is a list of lists: one inner list per query vector. That is why we write results[0].
  • Each hit has an id, a distance (here a cosine similarity, so higher is closer), and an entity holding the fields you asked for in output_fields.
  • limit is how many neighbours to return per query vector.

Without output_fields, you get only the id and distance. Ask for the fields you need and no more, because returning large fields slows a search.

Clean up

PYTHON
client.drop_collection(collection_name="demo_collection")

Dropping a collection deletes it and its data permanently, with no confirmation. Remember that on the day you are connected to a server that matters.

Try it
  1. Run the four calls above, but do not drop the collection yet.
  2. Change the query vector to point toward the tech entities, for example [0.05, 0.85, 0.2, 0.0], and search again.
  3. Set limit=4 and look at the order of all four results.
the tech sentences now come first, and the animal sentences come last. You have just done semantic search: the ranking changed because the direction of the query changed, with no keyword involved.

Designing a collection properly: schema and index

The quick setup has fixed field names and one vector field. The moment you want your own names, a text field with a declared length, or more control, you define a schema yourself. This is what real projects do.

Defining the schema

PYTHON
from pymilvus import MilvusClient, DataType

client = MilvusClient("./milvus_demo.db")

schema = MilvusClient.create_schema(
    auto_id=False,
    enable_dynamic_field=True,
)

schema.add_field(field_name="doc_id", datatype=DataType.INT64, is_primary=True)
schema.add_field(field_name="embedding", datatype=DataType.FLOAT_VECTOR, dim=64)
schema.add_field(field_name="title", datatype=DataType.VARCHAR, max_length=256)
schema.add_field(field_name="category", datatype=DataType.VARCHAR, max_length=64)
schema.add_field(field_name="year", datatype=DataType.INT64)

Walk through what each choice means.

auto_id=False says you supply the primary key yourself. Set it to True and Milvus generates ids, in which case you must not include the id field in your inserted data. Supplying your own ids is usually better, because then you can find, update, and delete a record using the identifier from your own system.

enable_dynamic_field=True keeps the flexibility of storing undeclared keys. It is convenient but it hides typos: if you insert "catgory" by mistake, Milvus quietly stores it as dynamic data instead of complaining. Teams that want strictness set it to False.

The data types you will use most are INT64 and INT32 for whole numbers, FLOAT and DOUBLE for decimals, BOOL, VARCHAR for strings, JSON for nested objects, and ARRAY for lists of one type. A VARCHAR must declare a max_length, and the limit is 65,535 bytes. Note bytes: Python's len() counts characters, and a single Arabic or emoji character takes several bytes, so text in Arabic reaches the limit sooner than the character count suggests. If you hit a length error, measure with len(s.encode("utf-8")).

Vector fields use FLOAT_VECTOR for the common 32-bit float embeddings, with dim set to your model's output size. Other vector types exist (16-bit float, 8-bit integer, binary, and sparse for keyword-style vectors), and they matter for efficiency and for full-text search, but FLOAT_VECTOR is the one to learn first. A collection may have more than one vector field, up to ten.

Defining the index

You tell Milvus how to index each field through an index-parameters object:

PYTHON
index_params = client.prepare_index_params()

index_params.add_index(
    field_name="embedding",
    index_type="AUTOINDEX",
    metric_type="COSINE",
)

client.create_collection(
    collection_name="articles",
    schema=schema,
    index_params=index_params,
)

Because you passed index_params, Milvus builds the index and loads the collection as part of creation. Check it:

PYTHON
print(client.get_load_state(collection_name="articles"))
print(client.describe_collection(collection_name="articles"))
TEXT
{'state': <LoadState: Loaded>}

describe_collection prints the schema back to you, which is the best way to confirm what you actually created. When something later does not match your expectations, this is the first call to make.

You can also index a scalar field to speed up filters on it. For a text field you filter by exact value, an INVERTED index is the general-purpose choice:

PYTHON
index_params.add_index(field_name="category", index_type="INVERTED")

Add scalar indexes only for fields you filter on frequently and on large collections. For a few thousand rows they change nothing you can measure.

The other way: create, then index, then load

For completeness, here is the explicit sequence, which you will meet in older material and which shows what the shortcut hides. Create the collection with only a schema, add the index separately, then load it:

PYTHON
client.create_collection(collection_name="articles2", schema=schema)
client.create_index(collection_name="articles2", index_params=index_params)
client.load_collection(collection_name="articles2")

If you create a collection with only a schema and skip the last two steps, it exists but is not loaded and not indexed, and your first search fails with collection not loaded or index not found. That is the bug behind most beginner questions, and now you know its cause.

You cannot change a vector's dimension later The dim of a vector field is fixed when the collection is created. If you switch embedding models and the new one outputs a different length, you create a new collection and re-insert everything. This is also why you should write the model name and dimension down somewhere (a collection property, a README, or the field name itself, such as embedding_384). Mixing vectors from two different models in one field is a silent disaster: Milvus will accept them if the lengths match, and your search results will simply be wrong.
Try it
  1. Create the articles collection with the schema above.
  2. Run describe_collection and find each of the five fields in the output.
  3. Try to insert an entity whose embedding has 63 numbers instead of 64, and read the error.
an error mentioning a dimension mismatch. Milvus validates every vector against the declared dim, which is a safety net you want. Fix the vector length and the insert succeeds.

Inserting, upserting, and deleting data

Data changes are straightforward, but there are a few behaviours worth knowing before they surprise you.

Insert in batches

PYTHON
rows = []
for i in range(1000):
    rows.append({
        "doc_id": i,
        "embedding": make_vector(i),    # your own function returning 64 floats
        "title": f"Article {i}",
        "category": "news" if i % 2 == 0 else "blog",
        "year": 2020 + i % 6,
    })

client.insert(collection_name="articles", data=rows)

Insert a list of dictionaries, not one row at a time. A call has overhead (a network round trip, validation, a write to the log), so inserting a thousand rows in one call is far faster than a thousand calls of one row. A reasonable beginner habit is batches of a few hundred to a few thousand rows. Very large single requests are rejected: the request size limit is 64 MB, so a million 768-dimensional vectors in one call will fail, and you should chunk them.

Inserting a primary key that already exists does not replace the old row. Primary key uniqueness is not enforced on insert, so you can end up with duplicates, which then show up twice in search results. If you might re-run an import, use upsert instead.

Upsert

PYTHON
client.upsert(
    collection_name="articles",
    data=[{"doc_id": 5, "embedding": make_vector(5), "title": "Article 5 (revised)",
           "category": "news", "year": 2026}],
)

Upsert means "insert, or replace the row if the primary key already exists". It is the safe default for any pipeline that might run twice. Milvus 3.0 also supports partial updates through upsert, where you send only the fields that changed, but as a beginner, send full rows and you will never be surprised.

Delete

You can delete by primary key or by a filter expression:

PYTHON
client.delete(collection_name="articles", ids=[0, 1, 2])
client.delete(collection_name="articles", filter="category == 'blog' and year < 2022")

Be careful with filter deletes: they remove every matching row, and a loose filter can wipe most of a collection. A good habit is to run the same expression through query first, check how many rows match, and only then delete.

Writes are not instantly visible everywhere

Milvus writes your data to a log first and processes it afterwards, so there can be a brief delay before a new row appears in search results. How much delay you tolerate is the consistency level. The four levels are Strong (always see the latest writes, at some latency cost), Bounded (the default, which allows a small lag), Session (you always see your own writes), and Eventually (fastest, no guarantee). You set it when you create the collection. For a beginner, the useful rule is: if you insert a row and immediately cannot find it, that is probably consistency, not a bug. Wait a moment, or create the collection with consistency_level="Strong" while learning.

Also note that deletes and upserts leave behind old data until a background process called compaction removes it. You do not run it by hand, and your searches never see deleted rows, but you may notice that storage does not shrink at once.

Try it
  1. Insert the 1,000 rows above (invent make_vector with random numbers if you like).
  2. Upsert doc_id 5 with a new title, then fetch it with client.get(collection_name="articles", ids=[5]).
  3. Delete ids 0 to 9 and fetch id 0 again.
the fetch of id 5 shows the revised title, once, not two copies. After the delete, fetching id 0 returns an empty list. Upsert replaced the row; delete removed it.

Searching: the everyday operations

Search is why you installed Milvus, and it has more depth than the first example showed. Three operations cover most beginner needs: vector search, search with a filter, and query.

Search with a filter

Real applications almost never want "the nearest thing to this vector, full stop". They want "the nearest article, but only from the news category and from 2024 onward". Milvus lets you attach a filter to a search. Milvus applies it as part of the search, so you get your full limit of matching results rather than a handful that survived filtering afterwards.

PYTHON
results = client.search(
    collection_name="articles",
    data=[make_vector(42)],
    limit=5,
    filter="category == 'news' and year >= 2024",
    output_fields=["title", "category", "year"],
)

for hit in results[0]:
    print(hit["id"], round(hit["distance"], 3), hit["entity"]["title"], hit["entity"]["year"])

The filter is a string written in Milvus's own expression language, which looks like a small piece of Python or SQL. The operators you need first:

  • Comparison: ==, !=, <, >, <=, >=
  • Ranges: "2020 < year < 2024"
  • Membership: "category in ['news', 'blog']" and not in
  • Text patterns: title like "Article 1%" (starts with), like "%2" (ends with), like "%news%" (contains)
  • Logic: and or &&, or or ||, not, and parentheses

Note the quoting: the expression is one Python string, and string values inside it are in a second set of quotes. Mixing them up (filter="category == news") makes Milvus look for a field called news, and the error complains about a field, not a value.

For values that come from users or from variables, use filter templating instead of gluing strings together, which both avoids quoting mistakes and avoids a user supplying a value that rewrites your expression:

PYTHON
results = client.search(
    collection_name="articles",
    data=[make_vector(42)],
    limit=5,
    filter="category == {cat} and year >= {min_year}",
    filter_params={"cat": "news", "min_year": 2024},
    output_fields=["title"],
)

The same filter_params works with query and delete. It is the equivalent of parameterised queries in SQL, and the same habit is worth having.

Query: no vectors involved

query retrieves rows by scalar conditions or by primary key. It does not rank by similarity, and it does not need a query vector:

PYTHON
rows = client.query(
    collection_name="articles",
    filter="category == 'blog' and year == 2025",
    output_fields=["title", "year"],
    limit=10,
)

same_rows = client.query(collection_name="articles", ids=[3, 7], output_fields=["title"])

Use query to inspect what is stored, to count and sample rows while debugging, and to look up records by id. When you only need ids, client.get(collection_name="articles", ids=[3, 7]) is the shortest form. Always pass a limit when querying with a filter: a filter that matches a million rows returns an enormous response otherwise, and Milvus caps the window of results you can request in one call at 16,384.

Searching several queries at once

Because data is a list, you can send many query vectors in a single call and get one inner list back for each:

PYTHON
batch = client.search(
    collection_name="articles",
    data=[make_vector(1), make_vector(2), make_vector(3)],
    limit=3,
    output_fields=["title"],
)
print(len(batch))        # 3 lists, one per query
print(len(batch[0]))     # up to 3 hits for the first query

Batching is much more efficient than three separate calls, and it is the right shape for any service that handles many users at once.

Why you sometimes get fewer results than limit

If you ask for ten results and receive six, check these in order. The collection might simply hold fewer than ten matching rows. A filter might be eliminating most of them. Duplicate primary keys (from repeated plain inserts) can collapse in the results. Or consistency has not yet exposed recent writes. A short result list is almost always one of these four, and almost never a defect in the search itself.

Debug with query before you debug with search When a search returns something odd, run a query with the same filter and count the matches. If the filter matches nothing, the problem is the filter or the data, and no amount of vector tuning would help. If it matches plenty, the problem is on the vector side. Separating the two halves saves hours.
Try it
  1. Search with no filter and note the top five ids.
  2. Repeat with filter="category == 'blog'" and compare the ids.
  3. Run a query with the same filter and a limit of 3, and print the rows.
every filtered hit is in the blog category, and you still receive five results because the filter was applied during the search. The query shows three plain rows with no distance field, because a query does not rank.

Seeing what you stored: the web interface and Attu

Command-line debugging is useful, but a graphical view of your collections makes many problems obvious in seconds. You already met the built-in WebUI at http://127.0.0.1:9091/webui/, which is useful for checking that the server is alive and for viewing basic information.

For a richer view, the Milvus project provides Attu, a graphical administration tool. You can browse collections, see schemas and indexes, run searches by pasting a vector, and inspect the data. The simplest way to run it is another container:

BASH
docker run -d --name attu -p 3000:3000 \
  -e MILVUS_ADDRESS=host.docker.internal:19530 \
  -v attu-data:/data \
  zilliz/attu:v3.0.0

Then open http://localhost:3000. The environment variable is called MILVUS_ADDRESS, not MILVUS_URL, and a wrong name does not fail loudly; Attu simply asks you for the address in its login page. The special hostname host.docker.internal lets a container reach a service on your host machine on Docker Desktop. On Linux it may not resolve, and you then use the host's IP address or put both containers on the same Docker network. Desktop applications of Attu also exist for macOS, Linux, and Windows if you prefer not to run a container. Attu 3.x works with Milvus 2.6 and 3.x, so match the tool to your server's generation.

Try it
  1. Run Attu (or open the WebUI) against your Docker standalone server.
  2. Find the articles collection and open its schema.
  3. Confirm the vector field's dimension, the index type, and the load state.
the three facts you set in code, displayed visually. If the load state says "not loaded", you have found the cause of a failing search without running a line of Python.

Configuration, security, and what to change first

Milvus works out of the box, which is good for learning and dangerous for anything shared. Here is what to know on day one.

Configuration lives in a YAML file, and most of it needs a restart

Milvus is configured by a large file, milvus.yaml, with sections for etcd, object storage, ports, limits, and more. You almost never edit the shipped file. Instead you override individual settings in a small file called user.yaml. With the embedded script, user.yaml sits in the folder where you ran it; in the Docker Compose setup you write it to /milvus/configs/user.yaml inside the container. After any change you must restart the container (docker restart milvus-standalone), because most Milvus settings are not reloaded while running. "I changed the config and nothing happened" is the single most common configuration complaint, and the answer is nearly always to restart.

Authentication is off by default

A fresh Milvus accepts any connection with no password. That is fine on a laptop and unacceptable on a server anyone can reach. The default user is root with the password Milvus, and these are publicly documented. To turn authentication on, set this in user.yaml and restart:

user.yaml
common:
  security:
    authorizationEnabled: true

Once it is on, clients must connect with a token such as root:Milvus, and you should immediately change the root password. Better still, set common.security.defaultRootPassword to a strong value before the first start so the documented default never exists on your server. Create separate users for applications through client.create_user and give them only the privileges they need; that role-based access control is covered in the Senior part. For now, remember the rule that matters: never expose port 19530 to the internet with the defaults.

Version 3.0.1 also raised the cost of password hashing, and the release notes say credentials need to be rotated after upgrading to it. If you ever upgrade an old installation, treat that as a reminder to change passwords.

The ports you will see

  • 19530: client connections (gRPC) and the RESTful API. This is the port your code uses.
  • 9091: the web interface, health, and Prometheus metrics.
  • 2379: etcd, inside the setup.
  • 9000 and 9001: MinIO, the object store in the Compose setup.

If a firewall or a cloud security group sits in front of your server, opening 19530 to the right clients is all an application needs. Keep 9091 and the storage ports closed to the outside.

Hardware, in one honest paragraph

Vector search is memory-hungry. Milvus loads your vectors and index into memory to answer queries, and as a rough guide the loaded size is the raw size of your vectors (number of vectors times dimension times four bytes for float32) plus index overhead. One million vectors of dimension 768 is about 3 GB of raw data. That is why the standalone minimum is 8 GB of RAM and why the dimension of your embedding model matters for cost. Smaller embedding models are not only faster; they are cheaper to host. Later parts cover ways to reduce memory, such as quantised indexes and memory-mapped files.

Try it
  1. On a throwaway Docker standalone, create user.yaml with the authorization setting above and restart the container.
  2. Connect with MilvusClient(uri="http://localhost:19530") and no token, then list collections.
  3. Connect again with token="root:Milvus".
the tokenless attempt fails with an authentication error (the message mentions that the user has not authenticated), and the attempt with the token works. Then change the password for the account you just secured.

Reading errors: the common ones and what they mean

Milvus errors reach Python as a MilvusException carrying a numeric code and a message. Read the message; it is usually precise. These are the ones a beginner meets.

collection not loaded. You searched or queried a collection that exists but is not in memory. Cause: it was created with only a schema, or you released it. Fix: make sure it has an index, then call client.load_collection("name"), and confirm with get_load_state. A related message, collection not fully loaded, means loading is still in progress; wait and retry.

collection not found. The name is wrong, or you are connected to a different database than the one holding it. Fix: print client.list_collections() and compare spelling and case carefully.

index not found. You tried to load or search a vector field that has no index. Fix: create the index first, then load.

invalid parameter and its relatives. A catch-all for bad inputs. The usual beginner causes are a vector of the wrong dimension, an unsupported combination of metric and index, or a limit above 16,384. Read the rest of the message for the field name and the expected value.

A length error such as the length (398324) of json field exceeds max length (65536). A string or JSON value is longer than its field allows, measured in bytes. Fix: shorten or split the text (chunking is standard practice for documents anyway), or raise max_length up to 65,535.

rate limit exceeded. You sent requests faster than the configured limits permit. Fix: slow the client and retry with a short pause.

service not ready or service unavailable, or a connection refused. The server is still starting, has stopped, or cannot reach etcd or object storage. Fix: docker ps, then docker logs milvus-standalone, and in Compose check that milvus-etcd and milvus-minio are healthy.

Illegal instruction. The container crashed because the CPU lacks the SIMD instructions Milvus needs, which happens on some virtual machines and emulated environments. Check with the lscpu command shown earlier.

ConnectionConfigException: Illegal uri: [example.db]. You used a Milvus Lite file path with a very old pymilvus. Upgrade pymilvus.

ModuleNotFoundError: No module named 'milvus_lite'. You are on Windows, where Milvus Lite is not available. Use WSL 2 or Docker.

The error is rarely where you first look Three beginner problems account for most "Milvus is broken" reports: the collection was not loaded, the vector length does not match the schema, and the server was not actually running. Check those three, in that order, before reading anything else. If a search finally runs but the answers look wrong, suspect the embeddings (a mismatched or changed model) before suspecting Milvus.
Try it
  1. Create a collection with only a schema (no index_params) and search it.
  2. Read the exception text.
  3. Create an index, call load_collection, and search again.
the first search raises an error about the missing index or the collection not being loaded; after the index and load, the same search returns results. You have reproduced, and cured, the most common Milvus problem on purpose.

Housekeeping commands you will use every week

Beyond searching, a handful of calls keep a Milvus instance understandable. Group them by what you are trying to do.

See what exists. client.list_collections() returns the names, client.has_collection("name") returns true or false, and client.describe_collection("name") returns the schema, the number of shards, the consistency level, and other properties. client.list_indexes("name") and client.describe_index(collection_name="name", index_name="...") show the indexes. Make describe_collection a reflex when you open a collection someone else created.

Free memory without losing data. client.release_collection("name") unloads a collection from memory. Its data and index remain on disk and in object storage; it simply cannot be searched until you call load_collection again. Releasing collections you are not using is the simplest way to stay within your memory budget on a small machine. client.get_load_state("name") tells you where you stand.

Rename and remove. client.rename_collection(old_name="a", new_name="b") renames a collection, and client.drop_collection("name") deletes it for good. Dropping an index requires releasing the collection first, which is a rule that surprises people: Milvus will not let you pull an index out from under a loaded collection.

Count what you have. A plain row count is a common need. Milvus 3.0 added query aggregation (functions such as count, sum, avg, min and max with grouping), described in the official query documentation. Check that page for the exact syntax of your client version. A simple fallback while learning is to run a query with a filter and a sensible limit and count the rows returned. Recent writes may take a moment to be visible, so for a precise number right after inserting, create the collection with Strong consistency.

Partitions, briefly. A collection can be divided into partitions, which are named subsets, and you can search only some of them. A new collection always has one called _default. Beginners rarely need more, and Milvus Lite does not support them at all. When you eventually need to separate data by customer or by tenant, a feature called the partition key does this more cleanly, and the Mid-level part covers it.

A safe rhythm for experiments Give each experiment its own collection with a name that says what it tests, such as exp_minilm_cosine. Collections are cheap. Dropping a whole collection and rebuilding it is far less error-prone than trying to repair one in place, and it mirrors how you will later re-index when you change embedding models.
Try it
  1. List your collections and describe one of them.
  2. Release the collection, then search it and read the error.
  3. Load it again and search successfully.
after release the search fails with the not-loaded error you now recognise, and after load it succeeds. Release and load are the on and off switch for searchability.

Putting it all together: a small semantic search project

Time to build something complete. We will make a tiny question-and-answer search over a handful of sentences, using only what you have learned. To keep the code runnable without downloading a model, we use a deliberately simple stand-in for an embedding model: it hashes each word into a slot of a 64-number vector and normalises the result. This is a toy. It matches shared words rather than meaning, so it will not find synonyms. Its job is to let you run the whole pipeline now. In a real project you replace one function with a call to a genuine embedding model and change nothing else except the dimension.

Step one: the embedding function and the data

PYTHON
import hashlib
import math

DIM = 64

def embed(text: str) -> list[float]:
    """Toy embedding: hash each word into one of DIM slots, then normalise."""
    vec = [0.0] * DIM
    for word in text.lower().split():
        word = word.strip(".,?!")
        slot = int(hashlib.md5(word.encode("utf-8")).hexdigest(), 16) % DIM
        vec[slot] += 1.0
    norm = math.sqrt(sum(x * x for x in vec)) or 1.0
    return [x / norm for x in vec]

docs = [
    {"id": 1, "text": "Reset your password from the account settings page.", "section": "account"},
    {"id": 2, "text": "Invoices are emailed on the first day of each month.", "section": "billing"},
    {"id": 3, "text": "You can change your billing address in the account settings page.", "section": "billing"},
    {"id": 4, "text": "Our support team answers email within one business day.", "section": "support"},
    {"id": 5, "text": "Delete your account from the account settings page.", "section": "account"},
]

Step two: build the collection

PYTHON
from pymilvus import MilvusClient, DataType

client = MilvusClient("./support_demo.db")   # or uri="http://localhost:19530"

if client.has_collection("support_docs"):
    client.drop_collection("support_docs")

schema = MilvusClient.create_schema(auto_id=False, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("vector", DataType.FLOAT_VECTOR, dim=DIM)
schema.add_field("text", DataType.VARCHAR, max_length=2000)
schema.add_field("section", DataType.VARCHAR, max_length=64)

index_params = client.prepare_index_params()
index_params.add_index(field_name="vector", index_type="AUTOINDEX", metric_type="COSINE")

client.create_collection(
    collection_name="support_docs",
    schema=schema,
    index_params=index_params,
)

We turned the dynamic field off so that a typo in a field name is an error rather than a silently stored extra. Because we passed index_params, the collection is indexed and loaded when the call returns.

PYTHON
rows = [{**d, "vector": embed(d["text"])} for d in docs]
client.insert(collection_name="support_docs", data=rows)

def ask(question: str, section: str | None = None, k: int = 2):
    kwargs = {}
    if section:
        kwargs["filter"] = "section == {s}"
        kwargs["filter_params"] = {"s": section}
    hits = client.search(
        collection_name="support_docs",
        data=[embed(question)],
        limit=k,
        output_fields=["text", "section"],
        **kwargs,
    )
    for hit in hits[0]:
        print(f'{hit["distance"]:.2f}  [{hit["entity"]["section"]}]  {hit["entity"]["text"]}')

ask("how do I change my account settings")
print("--- billing only ---")
ask("how do I change my account settings", section="billing")

You should see account and billing sentences about the settings page ranked first, because they share the words "account", "settings" and "page" with the question. The second call restricts the search to the billing section, so only billing rows can appear however well another section matches. If a hit's text is not what you expect, remember what the toy embedding does: it counts shared words. With a real embedding model the first query would also find "Reset your password" through meaning alone.

Step four: maintain the data

PYTHON
client.upsert(collection_name="support_docs", data=[{
    "id": 4,
    "vector": embed("Our support team answers email within four business hours."),
    "text": "Our support team answers email within four business hours.",
    "section": "support",
}])

client.delete(collection_name="support_docs", filter="section == 'billing'")

print(client.query(collection_name="support_docs", filter="id >= 0",
                   output_fields=["id", "section"], limit=10))

The upsert replaces row 4 by its primary key. The delete removes both billing rows. The final query lists what remains, which should be ids 1, 4 and 5.

Step five: swap in a real model

The only function that changes is embed. Choose an embedding model, set DIM to the length of its output, and create a new collection, since the dimension of an existing one cannot change. Keep the model name written down next to the collection. After that swap the search finds sentences by meaning, and the toy's limitations disappear. The retrieval step of a RAG system is exactly this code: embed the user's question, search, and pass the returned texts to a language model. Tools such as LangChain and LlamaIndex wrap these calls, and it helps to have written them once by hand.

Try it
  1. Run all five steps against Milvus Lite or Docker standalone.
  2. Ask a question that shares no words with any document and see how the toy embedding fails.
  3. Write down which single function you would replace to fix it, and what else must change.
the answer is the embed function, plus the DIM constant and a fresh collection. The database code stays identical, which is the point: Milvus is independent of how you produce vectors.

What you can now do, and what comes next

You can start Milvus in the mode that fits your situation, connect with MilvusClient, and explain the five nouns: collection, field, entity, index, and metric. You can create a collection in one call or define a custom schema with an index, and you know that loading is what makes it searchable. You can insert in batches, upsert safely, delete by id or filter, search with and without filters, query by scalar conditions, and read the most common errors without panic. You also know the two decisions that matter most before you scale: the embedding model fixes the dimension, and the defaults leave the server unauthenticated.

Where you will want to go next, in the Mid-level part: choosing and tuning an index such as HNSW or IVF rather than leaving it to AUTOINDEX, hybrid search that combines vector similarity with full-text keyword matching, partition keys for multi-tenant data, consistency levels in depth, schema changes on live collections, and server-side embedding functions. After that, the Senior part covers architecture, failure modes, scaling, security, upgrades, and operating Milvus as a platform. Neighbouring guides worth reading alongside: Qdrant, pgvector, and Chroma show how other vector stores answer the same questions, which is the best way to decide whether Milvus is the right fit; Ragas shows how to measure whether retrieval is actually good.

Sources