This is part one of three. It covers everything you need to do real work with Milvus as a beginner, not a teaser. By the end you can start a Milvus server on your laptop (or skip the server entirely and run it inside a Python script), design a collection, store embeddings with their metadata, search them by meaning, filter by scalar fields, update and delete data, read the common errors, and build a small semantic search project from scratch. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. Vector search feels abstract until you have watched your own query return the right sentence, and these ideas only stick once you have also watched a search fail because the collection was not loaded.
This guide targets Milvus 3.0.2, the current stable release at the time of writing, with the matching Python client, pymilvus 3.0.2. A great deal of what you find online was written for Milvus 2.5 or earlier. We flag the places where that older material is now wrong.
What Milvus is, and the problem it solves
Milvus is an open-source vector database. It stores lists of numbers called vectors (also called embeddings), together with ordinary fields such as a title or a date, and it answers one question very quickly: which stored vectors are closest to this one?
To see why that matters, start with what an embedding is. A machine learning model can turn a piece of text, an image, or a sound clip into a list of a few hundred or a few thousand numbers. The model is trained so that things with similar meaning end up with similar numbers. The sentence "How do I reset my password?" and the sentence "I forgot my login credentials" share almost no words, yet an embedding model places them close together. "The quarterly revenue grew" lands far away from both. Closeness in number space stands in for closeness in meaning.
That gives you a new kind of search. A traditional database finds rows where a column equals a value, or where a text field contains a word. A vector database finds rows whose meaning is near the meaning of your query, even when no word matches. This is the engine behind semantic search, recommendation, duplicate detection, image search, and the retrieval step of retrieval-augmented generation (RAG), where a language model is given relevant passages pulled from your own documents before it answers.
Note what Milvus does and does not do. Milvus does not create embeddings for you in the simple case. An embedding model does that (a model you run, or a hosted service), and you hand the resulting vectors to Milvus. Milvus stores them, builds an index over them, and searches them. Newer versions can also call an embedding provider from the server side, which is a mid-level topic. For now, the mental split is: the model understands meaning, Milvus finds neighbours.
What came before
You could always do this with brute force: keep every vector in a NumPy array, compute the distance from your query to each one, sort, and take the top few. For ten thousand vectors that is fine and you should not feel bad about it. At ten million vectors it is too slow, and it also leaves you to build everything around it: persistence, metadata filtering, updates and deletes, access from several services, backups.
Libraries such as FAISS solved the speed part by providing fast approximate-nearest-neighbour indexes, but a library is not a database. It lives inside your process and leaves storage, filtering, and concurrent access to you. Milvus is the database layer around that idea: a server (or an embedded library) that persists data, indexes it, filters it, scales, and speaks a network protocol that many programs can use. There is a dedicated guide to FAISS if you want to see the library underneath that world.
The three ways to run Milvus
Milvus comes in three deployment modes, and picking the right one is your first decision.
- Milvus Lite is a Python library. You install it with pip, point it at a file such as
./milvus_demo.db, and it runs inside your Python process. There is no server and no Docker. It is for learning and prototyping, and it runs on Linux and macOS but not on Windows. - Standalone is a single Milvus process in one container, plus two companions it depends on (etcd for metadata and an object store for data files). It is the right choice for a laptop, a small server, or a first real project.
- Distributed runs Milvus as several cooperating services on Kubernetes. It is for large data and high traffic, and it is covered in the Mid-level and Senior parts.
There is also Zilliz Cloud, the managed service run by the company behind Milvus, which gives you the same API without operating anything yourself. The code you write in this guide works against all of these with a one-line change to the connection address, which is a real advantage when you start small and grow.
- Write three sentences about different topics, for example a cooking tip, a football result, and a database tip.
- Write a fourth sentence that is related to one of them but shares no important words with it.
- Decide by hand which of the first three the fourth is closest to.
The core ideas: five nouns
Milvus has a larger vocabulary than most beginner tools, but five nouns carry nearly all of it. Learn these and the documentation becomes readable.
Collection. A collection is the unit you work with, and it is the closest thing to a table. You create one, insert into it, search it, and drop it. A collection has a schema that says what fields every record contains. A database is just a namespace that groups collections. There is always one called default, and a beginner can work in it forever.
Field and schema. A schema lists the fields. Every collection needs exactly one primary key field (an integer or a string) that uniquely identifies each record, and at least one vector field that holds the embedding. It can also have any number of scalar fields, which is Milvus's word for ordinary values: text, numbers, booleans, JSON, arrays. The vector field has a dimension, the length of the list, and every vector you store in that field must have exactly that length. If your embedding model produces 768 numbers, the field has dimension 768, and a vector of 767 numbers is rejected.
Entity. An entity is one record: a row. In Python you write it as a dictionary, for example {"id": 1, "vector": [...], "text": "..."}.
Index. An index is a data structure built over a vector field so that searches do not have to compare against everything. Without an index Milvus can only do an exact comparison with every vector (a brute-force scan). With one it does an approximate search: dramatically faster, with a small, tunable chance of missing a true neighbour. Approximate nearest-neighbour search is the central trade-off of the whole field, and we return to it below.
Metric type. To find "closest", Milvus needs a definition of distance. For ordinary float vectors the three metrics are COSINE (the angle between vectors, ignoring their length), L2 (straight-line Euclidean distance), and IP (inner product). The rule is to use the metric your embedding model was trained for. For most text embedding models that is cosine similarity, and cosine is also the default for the quick setup you will use in a moment.
L2 a smaller number means closer, so the best hit has the smallest distance. With COSINE and IP a larger number means more similar, so the best hit has the largest. Milvus handles the ordering for you and returns the best results first, but the field it returns is called distance in every case. If you print a score and it looks "backwards", check your metric before assuming a bug.
Load and release, the idea that surprises everyone
Here is the concept that trips up almost every newcomer. For a search to work, the collection must be loaded. Loading means the data and its index are brought into memory on the server so that queries can be answered fast. A collection that exists but is not loaded can be inserted into, but cannot be searched, and you get the error collection not loaded.
The reassuring part: when you create a collection with the quick setup (shown below), or when you pass index parameters at creation time, Milvus loads it for you. You meet load_collection only when you take a different route, or after you release a collection to free memory. Remember the pairing, though, because it is the first thing to check when a search fails.
Segments, in one paragraph
You do not need to manage them, but you will see the word in logs. Milvus stores incoming data in segments. New writes land in a growing segment. When a segment is big enough or you ask for a flush, it is sealed, turned into immutable files in object storage, and indexed. Searches look across both. This is why a freshly inserted row can occasionally be invisible for a moment and why the consistency level, discussed later, exists.
- Imagine a collection of support articles. Write down its primary key, its vector field (with a dimension of 384), and three scalar fields.
- For each scalar field, decide whether you would ever filter by it.
- Say which metric you would choose if your model's documentation says it was trained with cosine similarity.
id (integer key), vector (dimension 384), title (text), category (text), year (integer), with the metric set to COSINE. Fields you filter on belong in the schema; they are what makes a search "find similar articles, but only from 2025".
Approximate nearest-neighbour search, in plain words
Before touching code, spend two minutes on the idea that explains almost every setting you will meet later, because it also explains the interview questions.
Searching a million vectors exactly means a million distance calculations per query. Approximate search avoids most of them by organising the vectors in advance. Two families dominate. Graph indexes (the best known is HNSW) link each vector to some of its near neighbours and then walk that graph from a starting point toward the query, always stepping to a closer vertex. Cluster indexes (the IVF family) split the space into regions around centre points, then search only the few regions nearest the query. Both examine a small fraction of the data and return answers that are almost always the true nearest ones.
The word "almost" has a name: recall, the fraction of true nearest neighbours that the search actually returns. Higher recall costs more time and memory. Each index has parameters that move you along that line, and the defaults are sensible. As a beginner you will not tune them. You will let Milvus choose with the setting called AUTOINDEX, which picks an index and parameters for you. That is what the quick setup does.
Keep two facts. First, an index is built over a vector field and you must have one before searching at scale. Second, FLAT is the one index that is exact: it compares against everything and never misses. It is the right choice for small datasets and for checking what the approximate indexes are missing. Milvus Lite, in fact, always uses FLAT regardless of what you ask for, which is fine because it is meant for small data.
Installing and checking the setup
You have three options. Start with Milvus Lite if you only want to learn the API, because it needs nothing but Python. Move to Docker standalone when you want the real server, the visual tools, and a setup that resembles production.
Option one: Milvus Lite (Linux and macOS)
pip install -U "pymilvus[milvus-lite]"
That single command installs the Python client and the embedded engine. It supports Ubuntu 20.04 and later (x86_64 and arm64) and macOS 11 and later (Apple Silicon and Intel). It does not run on Windows. If you try, you get ModuleNotFoundError: No module named 'milvus_lite'. Windows users should use WSL 2 (a Linux environment inside Windows) or Docker, both described below.
Connecting is one line, and the "address" is just a file path:
from pymilvus import MilvusClient
client = MilvusClient("./milvus_demo.db")
print(client.list_collections())
[]
An empty list means the connection works and nothing is stored yet. The file milvus_demo.db appears next to your script and holds your data between runs. Milvus Lite has real limits: it does not support partitions, users and roles, or aliases, and it ignores some collection options. It is a learning tool, and a pleasant one. Later, milvus-lite dump can export a collection from the file so you can load it into a full Milvus.
Option two: Milvus standalone in Docker
Standalone is the real server. You need a machine with at least 8 GB of memory available (16 GB is recommended) and a CPU with SIMD instructions (SSE4.2, AVX, AVX2, or AVX-512), which any recent laptop has. On Linux you can check with:
lscpu | grep -e sse4_2 -e avx -e avx2 -e avx512
On a Mac, Docker Desktop's virtual machine should be given at least 2 virtual CPUs and 8 GB of memory, which you set in its settings. On Windows, use Docker Desktop with WSL 2 and keep your data in the Linux filesystem rather than the Windows one. If you have never used Docker, the Docker guide in this series covers containers from zero, and this guide assumes only that docker runs on your machine.
The quickest start is the official script, which launches a single container with an embedded etcd and local storage:
curl -sfL https://raw.githubusercontent.com/milvus-io/milvus/master/scripts/standalone_embed.sh -o standalone_embed.sh
bash standalone_embed.sh start
The script creates a container named milvus-standalone. Milvus listens for client connections on port 19530, and a small web interface and the metrics endpoint listen on port 9091. Data is stored in a volumes/milvus folder beside the script. The same script accepts stop, restart, delete, and upgrade. On Windows there is an equivalent standalone_embed.bat, run from PowerShell with Docker Desktop started as administrator.
The other official route is Docker Compose, which runs three containers: Milvus, plus separate etcd and MinIO (an S3-compatible object store) containers. It is closer to production and shows you the moving parts:
wget https://github.com/milvus-io/milvus/releases/download/v3.0.2/milvus-standalone-docker-compose.yml -O docker-compose.yml
sudo docker compose up -d
sudo docker compose ps
docker compose ps should list milvus-standalone, milvus-minio, and milvus-etcd as running (healthy). The file name differs slightly between documentation pages (.yaml on one, .yml on another), so if the download fails, check the release assets for the exact name. To tear everything down and delete the data, run sudo docker compose down and then sudo rm -rf volumes. Be careful: that deletes your stored vectors.
http://127.0.0.1:9091/webui/ in a browser. If a page loads, the server is up and you have ruled out a whole category of connection problems before your Python script ever runs. The same port serves Prometheus metrics at /metrics, which you will care about later.
Installing the Python client
pip install -U pymilvus
python -c "import pymilvus; print(pymilvus.__version__)"
For Milvus 3.0.2 you want pymilvus 3.0.2. Use a virtual environment so the client does not collide with other projects. There is also an optional extra, pip install "pymilvus[model]", that adds helper classes for generating embeddings. We will not need it, but you will see it in tutorials.
Now connect to your standalone server and prove the whole chain works:
from pymilvus import MilvusClient
client = MilvusClient(uri="http://localhost:19530", token="root:Milvus")
print(client.list_collections())
[]
The token here is user:password. A fresh Milvus installation has authentication turned off, so the token is ignored, but the documented default account is user root with password Milvus, and including it makes your code ready for the day authentication is on. We return to the security consequences in a later section.
from pymilvus import connections, Collection and calls such as connections.connect() and Collection(...). That older object-oriented style still exists in places, but the current documentation is built around MilvusClient, and so is this guide. Also, any architecture diagram that shows separate RootCoord, QueryCoord, DataCoord, or IndexNode services is from before Milvus 2.6, which merged and removed them. Do not copy commands that scale an "index node".
- Pick Milvus Lite (if you are on Linux or macOS and just want the API) or Docker standalone.
- Install pymilvus and connect with
MilvusClient. - Call
client.list_collections(). - If you used Docker, also open
http://127.0.0.1:9091/webui/.
[] and, for Docker, a loading web page. If the connection hangs or is refused, the container is not running yet (check docker ps) or is still starting; give it a minute and look at docker logs milvus-standalone.
Your first collection in four calls
The fastest path from nothing to a working vector search is the quick setup. In this section we use tiny three-number vectors so you can read every value; the next sections use realistic ones.
Create the collection
from pymilvus import MilvusClient
client = MilvusClient("./milvus_demo.db") # or uri="http://localhost:19530"
if client.has_collection("demo_collection"):
client.drop_collection("demo_collection")
client.create_collection(
collection_name="demo_collection",
dimension=4,
)
Two arguments are all it takes. Behind that short call, Milvus does a surprising amount, and you should know what, because you did not ask for it:
- It creates a primary key field named
id(an integer) and a vector field namedvectorwith dimension 4. - It turns on the dynamic field, which means you can insert extra keys that are not in the schema, and they are stored in a hidden JSON field.
- It builds an index on the vector field using AUTOINDEX, with the default metric
COSINE. - It loads the collection, so it is searchable immediately.
The quick setup is excellent for learning and for simple projects. Its cost is flexibility: field names are fixed, and the primary key is not auto-generated unless you pass auto_id=True. The custom schema in a later section removes those limits.
Insert some data
data = [
{"id": 0, "vector": [0.9, 0.1, 0.0, 0.1], "text": "Cats purr when they are content.", "topic": "animals"},
{"id": 1, "vector": [0.8, 0.2, 0.1, 0.0], "text": "Dogs need daily walks.", "topic": "animals"},
{"id": 2, "vector": [0.1, 0.9, 0.2, 0.0], "text": "Python is a programming language.", "topic": "tech"},
{"id": 3, "vector": [0.0, 0.8, 0.3, 0.1], "text": "Milvus stores vectors.", "topic": "tech"},
]
result = client.insert(collection_name="demo_collection", data=data)
print(result)
{'insert_count': 4, 'ids': [0, 1, 2, 3], 'cost': 0}
Each entity is a dictionary. The id and vector keys match the schema. The text and topic keys are not in the schema at all; they were accepted because the dynamic field is on, and you can search and filter on them just like declared fields. The return value tells you how many rows were inserted and which ids they received.
Search
query_vector = [0.85, 0.15, 0.05, 0.05] # a "cat-like" direction
results = client.search(
collection_name="demo_collection",
data=[query_vector],
limit=2,
output_fields=["text", "topic"],
)
for hit in results[0]:
print(hit["id"], round(hit["distance"], 3), hit["entity"])
0 0.999 {'text': 'Cats purr when they are content.', 'topic': 'animals'}
1 0.995 {'text': 'Dogs need daily walks.', 'topic': 'animals'}
Your exact decimals may differ slightly, but the shape is what matters. Read it closely, because it is the most common thing beginners misread:
datais a list of query vectors, even when you have one. You can search several vectors in one call.- The result is a list of lists: one inner list per query vector. That is why we write
results[0]. - Each hit has an
id, adistance(here a cosine similarity, so higher is closer), and anentityholding the fields you asked for inoutput_fields. limitis how many neighbours to return per query vector.
Without output_fields, you get only the id and distance. Ask for the fields you need and no more, because returning large fields slows a search.
Clean up
client.drop_collection(collection_name="demo_collection")
Dropping a collection deletes it and its data permanently, with no confirmation. Remember that on the day you are connected to a server that matters.
- Run the four calls above, but do not drop the collection yet.
- Change the query vector to point toward the tech entities, for example
[0.05, 0.85, 0.2, 0.0], and search again. - Set
limit=4and look at the order of all four results.
Designing a collection properly: schema and index
The quick setup has fixed field names and one vector field. The moment you want your own names, a text field with a declared length, or more control, you define a schema yourself. This is what real projects do.
Defining the schema
from pymilvus import MilvusClient, DataType
client = MilvusClient("./milvus_demo.db")
schema = MilvusClient.create_schema(
auto_id=False,
enable_dynamic_field=True,
)
schema.add_field(field_name="doc_id", datatype=DataType.INT64, is_primary=True)
schema.add_field(field_name="embedding", datatype=DataType.FLOAT_VECTOR, dim=64)
schema.add_field(field_name="title", datatype=DataType.VARCHAR, max_length=256)
schema.add_field(field_name="category", datatype=DataType.VARCHAR, max_length=64)
schema.add_field(field_name="year", datatype=DataType.INT64)
Walk through what each choice means.
auto_id=False says you supply the primary key yourself. Set it to True and Milvus generates ids, in which case you must not include the id field in your inserted data. Supplying your own ids is usually better, because then you can find, update, and delete a record using the identifier from your own system.
enable_dynamic_field=True keeps the flexibility of storing undeclared keys. It is convenient but it hides typos: if you insert "catgory" by mistake, Milvus quietly stores it as dynamic data instead of complaining. Teams that want strictness set it to False.
The data types you will use most are INT64 and INT32 for whole numbers, FLOAT and DOUBLE for decimals, BOOL, VARCHAR for strings, JSON for nested objects, and ARRAY for lists of one type. A VARCHAR must declare a max_length, and the limit is 65,535 bytes. Note bytes: Python's len() counts characters, and a single Arabic or emoji character takes several bytes, so text in Arabic reaches the limit sooner than the character count suggests. If you hit a length error, measure with len(s.encode("utf-8")).
Vector fields use FLOAT_VECTOR for the common 32-bit float embeddings, with dim set to your model's output size. Other vector types exist (16-bit float, 8-bit integer, binary, and sparse for keyword-style vectors), and they matter for efficiency and for full-text search, but FLOAT_VECTOR is the one to learn first. A collection may have more than one vector field, up to ten.
Defining the index
You tell Milvus how to index each field through an index-parameters object:
index_params = client.prepare_index_params()
index_params.add_index(
field_name="embedding",
index_type="AUTOINDEX",
metric_type="COSINE",
)
client.create_collection(
collection_name="articles",
schema=schema,
index_params=index_params,
)
Because you passed index_params, Milvus builds the index and loads the collection as part of creation. Check it:
print(client.get_load_state(collection_name="articles"))
print(client.describe_collection(collection_name="articles"))
{'state': <LoadState: Loaded>}
describe_collection prints the schema back to you, which is the best way to confirm what you actually created. When something later does not match your expectations, this is the first call to make.
You can also index a scalar field to speed up filters on it. For a text field you filter by exact value, an INVERTED index is the general-purpose choice:
index_params.add_index(field_name="category", index_type="INVERTED")
Add scalar indexes only for fields you filter on frequently and on large collections. For a few thousand rows they change nothing you can measure.
The other way: create, then index, then load
For completeness, here is the explicit sequence, which you will meet in older material and which shows what the shortcut hides. Create the collection with only a schema, add the index separately, then load it:
client.create_collection(collection_name="articles2", schema=schema)
client.create_index(collection_name="articles2", index_params=index_params)
client.load_collection(collection_name="articles2")
If you create a collection with only a schema and skip the last two steps, it exists but is not loaded and not indexed, and your first search fails with collection not loaded or index not found. That is the bug behind most beginner questions, and now you know its cause.
dim of a vector field is fixed when the collection is created. If you switch embedding models and the new one outputs a different length, you create a new collection and re-insert everything. This is also why you should write the model name and dimension down somewhere (a collection property, a README, or the field name itself, such as embedding_384). Mixing vectors from two different models in one field is a silent disaster: Milvus will accept them if the lengths match, and your search results will simply be wrong.
- Create the
articlescollection with the schema above. - Run
describe_collectionand find each of the five fields in the output. - Try to insert an entity whose
embeddinghas 63 numbers instead of 64, and read the error.
dim, which is a safety net you want. Fix the vector length and the insert succeeds.
Inserting, upserting, and deleting data
Data changes are straightforward, but there are a few behaviours worth knowing before they surprise you.
Insert in batches
rows = []
for i in range(1000):
rows.append({
"doc_id": i,
"embedding": make_vector(i), # your own function returning 64 floats
"title": f"Article {i}",
"category": "news" if i % 2 == 0 else "blog",
"year": 2020 + i % 6,
})
client.insert(collection_name="articles", data=rows)
Insert a list of dictionaries, not one row at a time. A call has overhead (a network round trip, validation, a write to the log), so inserting a thousand rows in one call is far faster than a thousand calls of one row. A reasonable beginner habit is batches of a few hundred to a few thousand rows. Very large single requests are rejected: the request size limit is 64 MB, so a million 768-dimensional vectors in one call will fail, and you should chunk them.
Inserting a primary key that already exists does not replace the old row. Primary key uniqueness is not enforced on insert, so you can end up with duplicates, which then show up twice in search results. If you might re-run an import, use upsert instead.
Upsert
client.upsert(
collection_name="articles",
data=[{"doc_id": 5, "embedding": make_vector(5), "title": "Article 5 (revised)",
"category": "news", "year": 2026}],
)
Upsert means "insert, or replace the row if the primary key already exists". It is the safe default for any pipeline that might run twice. Milvus 3.0 also supports partial updates through upsert, where you send only the fields that changed, but as a beginner, send full rows and you will never be surprised.
Delete
You can delete by primary key or by a filter expression:
client.delete(collection_name="articles", ids=[0, 1, 2])
client.delete(collection_name="articles", filter="category == 'blog' and year < 2022")
Be careful with filter deletes: they remove every matching row, and a loose filter can wipe most of a collection. A good habit is to run the same expression through query first, check how many rows match, and only then delete.
Writes are not instantly visible everywhere
Milvus writes your data to a log first and processes it afterwards, so there can be a brief delay before a new row appears in search results. How much delay you tolerate is the consistency level. The four levels are Strong (always see the latest writes, at some latency cost), Bounded (the default, which allows a small lag), Session (you always see your own writes), and Eventually (fastest, no guarantee). You set it when you create the collection. For a beginner, the useful rule is: if you insert a row and immediately cannot find it, that is probably consistency, not a bug. Wait a moment, or create the collection with consistency_level="Strong" while learning.
Also note that deletes and upserts leave behind old data until a background process called compaction removes it. You do not run it by hand, and your searches never see deleted rows, but you may notice that storage does not shrink at once.
- Insert the 1,000 rows above (invent
make_vectorwith random numbers if you like). - Upsert doc_id 5 with a new title, then fetch it with
client.get(collection_name="articles", ids=[5]). - Delete ids 0 to 9 and fetch id 0 again.
Searching: the everyday operations
Search is why you installed Milvus, and it has more depth than the first example showed. Three operations cover most beginner needs: vector search, search with a filter, and query.
Search with a filter
Real applications almost never want "the nearest thing to this vector, full stop". They want "the nearest article, but only from the news category and from 2024 onward". Milvus lets you attach a filter to a search. Milvus applies it as part of the search, so you get your full limit of matching results rather than a handful that survived filtering afterwards.
results = client.search(
collection_name="articles",
data=[make_vector(42)],
limit=5,
filter="category == 'news' and year >= 2024",
output_fields=["title", "category", "year"],
)
for hit in results[0]:
print(hit["id"], round(hit["distance"], 3), hit["entity"]["title"], hit["entity"]["year"])
The filter is a string written in Milvus's own expression language, which looks like a small piece of Python or SQL. The operators you need first:
- Comparison:
==,!=,<,>,<=,>= - Ranges:
"2020 < year < 2024" - Membership:
"category in ['news', 'blog']"andnot in - Text patterns:
title like "Article 1%"(starts with),like "%2"(ends with),like "%news%"(contains) - Logic:
andor&&,oror||,not, and parentheses
Note the quoting: the expression is one Python string, and string values inside it are in a second set of quotes. Mixing them up (filter="category == news") makes Milvus look for a field called news, and the error complains about a field, not a value.
For values that come from users or from variables, use filter templating instead of gluing strings together, which both avoids quoting mistakes and avoids a user supplying a value that rewrites your expression:
results = client.search(
collection_name="articles",
data=[make_vector(42)],
limit=5,
filter="category == {cat} and year >= {min_year}",
filter_params={"cat": "news", "min_year": 2024},
output_fields=["title"],
)
The same filter_params works with query and delete. It is the equivalent of parameterised queries in SQL, and the same habit is worth having.
Query: no vectors involved
query retrieves rows by scalar conditions or by primary key. It does not rank by similarity, and it does not need a query vector:
rows = client.query(
collection_name="articles",
filter="category == 'blog' and year == 2025",
output_fields=["title", "year"],
limit=10,
)
same_rows = client.query(collection_name="articles", ids=[3, 7], output_fields=["title"])
Use query to inspect what is stored, to count and sample rows while debugging, and to look up records by id. When you only need ids, client.get(collection_name="articles", ids=[3, 7]) is the shortest form. Always pass a limit when querying with a filter: a filter that matches a million rows returns an enormous response otherwise, and Milvus caps the window of results you can request in one call at 16,384.
Searching several queries at once
Because data is a list, you can send many query vectors in a single call and get one inner list back for each:
batch = client.search(
collection_name="articles",
data=[make_vector(1), make_vector(2), make_vector(3)],
limit=3,
output_fields=["title"],
)
print(len(batch)) # 3 lists, one per query
print(len(batch[0])) # up to 3 hits for the first query
Batching is much more efficient than three separate calls, and it is the right shape for any service that handles many users at once.
Why you sometimes get fewer results than limit
If you ask for ten results and receive six, check these in order. The collection might simply hold fewer than ten matching rows. A filter might be eliminating most of them. Duplicate primary keys (from repeated plain inserts) can collapse in the results. Or consistency has not yet exposed recent writes. A short result list is almost always one of these four, and almost never a defect in the search itself.
query with the same filter and count the matches. If the filter matches nothing, the problem is the filter or the data, and no amount of vector tuning would help. If it matches plenty, the problem is on the vector side. Separating the two halves saves hours.
- Search with no filter and note the top five ids.
- Repeat with
filter="category == 'blog'"and compare the ids. - Run a
querywith the same filter and a limit of 3, and print the rows.
distance field, because a query does not rank.
Seeing what you stored: the web interface and Attu
Command-line debugging is useful, but a graphical view of your collections makes many problems obvious in seconds. You already met the built-in WebUI at http://127.0.0.1:9091/webui/, which is useful for checking that the server is alive and for viewing basic information.
For a richer view, the Milvus project provides Attu, a graphical administration tool. You can browse collections, see schemas and indexes, run searches by pasting a vector, and inspect the data. The simplest way to run it is another container:
docker run -d --name attu -p 3000:3000 \
-e MILVUS_ADDRESS=host.docker.internal:19530 \
-v attu-data:/data \
zilliz/attu:v3.0.0
Then open http://localhost:3000. The environment variable is called MILVUS_ADDRESS, not MILVUS_URL, and a wrong name does not fail loudly; Attu simply asks you for the address in its login page. The special hostname host.docker.internal lets a container reach a service on your host machine on Docker Desktop. On Linux it may not resolve, and you then use the host's IP address or put both containers on the same Docker network. Desktop applications of Attu also exist for macOS, Linux, and Windows if you prefer not to run a container. Attu 3.x works with Milvus 2.6 and 3.x, so match the tool to your server's generation.
- Run Attu (or open the WebUI) against your Docker standalone server.
- Find the
articlescollection and open its schema. - Confirm the vector field's dimension, the index type, and the load state.
Configuration, security, and what to change first
Milvus works out of the box, which is good for learning and dangerous for anything shared. Here is what to know on day one.
Configuration lives in a YAML file, and most of it needs a restart
Milvus is configured by a large file, milvus.yaml, with sections for etcd, object storage, ports, limits, and more. You almost never edit the shipped file. Instead you override individual settings in a small file called user.yaml. With the embedded script, user.yaml sits in the folder where you ran it; in the Docker Compose setup you write it to /milvus/configs/user.yaml inside the container. After any change you must restart the container (docker restart milvus-standalone), because most Milvus settings are not reloaded while running. "I changed the config and nothing happened" is the single most common configuration complaint, and the answer is nearly always to restart.
Authentication is off by default
A fresh Milvus accepts any connection with no password. That is fine on a laptop and unacceptable on a server anyone can reach. The default user is root with the password Milvus, and these are publicly documented. To turn authentication on, set this in user.yaml and restart:
common:
security:
authorizationEnabled: true
Once it is on, clients must connect with a token such as root:Milvus, and you should immediately change the root password. Better still, set common.security.defaultRootPassword to a strong value before the first start so the documented default never exists on your server. Create separate users for applications through client.create_user and give them only the privileges they need; that role-based access control is covered in the Senior part. For now, remember the rule that matters: never expose port 19530 to the internet with the defaults.
Version 3.0.1 also raised the cost of password hashing, and the release notes say credentials need to be rotated after upgrading to it. If you ever upgrade an old installation, treat that as a reminder to change passwords.
The ports you will see
- 19530: client connections (gRPC) and the RESTful API. This is the port your code uses.
- 9091: the web interface, health, and Prometheus metrics.
- 2379: etcd, inside the setup.
- 9000 and 9001: MinIO, the object store in the Compose setup.
If a firewall or a cloud security group sits in front of your server, opening 19530 to the right clients is all an application needs. Keep 9091 and the storage ports closed to the outside.
Hardware, in one honest paragraph
Vector search is memory-hungry. Milvus loads your vectors and index into memory to answer queries, and as a rough guide the loaded size is the raw size of your vectors (number of vectors times dimension times four bytes for float32) plus index overhead. One million vectors of dimension 768 is about 3 GB of raw data. That is why the standalone minimum is 8 GB of RAM and why the dimension of your embedding model matters for cost. Smaller embedding models are not only faster; they are cheaper to host. Later parts cover ways to reduce memory, such as quantised indexes and memory-mapped files.
- On a throwaway Docker standalone, create
user.yamlwith the authorization setting above and restart the container. - Connect with
MilvusClient(uri="http://localhost:19530")and no token, then list collections. - Connect again with
token="root:Milvus".
Reading errors: the common ones and what they mean
Milvus errors reach Python as a MilvusException carrying a numeric code and a message. Read the message; it is usually precise. These are the ones a beginner meets.
collection not loaded. You searched or queried a collection that exists but is not in memory. Cause: it was created with only a schema, or you released it. Fix: make sure it has an index, then call client.load_collection("name"), and confirm with get_load_state. A related message, collection not fully loaded, means loading is still in progress; wait and retry.
collection not found. The name is wrong, or you are connected to a different database than the one holding it. Fix: print client.list_collections() and compare spelling and case carefully.
index not found. You tried to load or search a vector field that has no index. Fix: create the index first, then load.
invalid parameter and its relatives. A catch-all for bad inputs. The usual beginner causes are a vector of the wrong dimension, an unsupported combination of metric and index, or a limit above 16,384. Read the rest of the message for the field name and the expected value.
A length error such as the length (398324) of json field exceeds max length (65536). A string or JSON value is longer than its field allows, measured in bytes. Fix: shorten or split the text (chunking is standard practice for documents anyway), or raise max_length up to 65,535.
rate limit exceeded. You sent requests faster than the configured limits permit. Fix: slow the client and retry with a short pause.
service not ready or service unavailable, or a connection refused. The server is still starting, has stopped, or cannot reach etcd or object storage. Fix: docker ps, then docker logs milvus-standalone, and in Compose check that milvus-etcd and milvus-minio are healthy.
Illegal instruction. The container crashed because the CPU lacks the SIMD instructions Milvus needs, which happens on some virtual machines and emulated environments. Check with the lscpu command shown earlier.
ConnectionConfigException: Illegal uri: [example.db]. You used a Milvus Lite file path with a very old pymilvus. Upgrade pymilvus.
ModuleNotFoundError: No module named 'milvus_lite'. You are on Windows, where Milvus Lite is not available. Use WSL 2 or Docker.
- Create a collection with only a schema (no
index_params) and search it. - Read the exception text.
- Create an index, call
load_collection, and search again.
Housekeeping commands you will use every week
Beyond searching, a handful of calls keep a Milvus instance understandable. Group them by what you are trying to do.
See what exists. client.list_collections() returns the names, client.has_collection("name") returns true or false, and client.describe_collection("name") returns the schema, the number of shards, the consistency level, and other properties. client.list_indexes("name") and client.describe_index(collection_name="name", index_name="...") show the indexes. Make describe_collection a reflex when you open a collection someone else created.
Free memory without losing data. client.release_collection("name") unloads a collection from memory. Its data and index remain on disk and in object storage; it simply cannot be searched until you call load_collection again. Releasing collections you are not using is the simplest way to stay within your memory budget on a small machine. client.get_load_state("name") tells you where you stand.
Rename and remove. client.rename_collection(old_name="a", new_name="b") renames a collection, and client.drop_collection("name") deletes it for good. Dropping an index requires releasing the collection first, which is a rule that surprises people: Milvus will not let you pull an index out from under a loaded collection.
Count what you have. A plain row count is a common need. Milvus 3.0 added query aggregation (functions such as count, sum, avg, min and max with grouping), described in the official query documentation. Check that page for the exact syntax of your client version. A simple fallback while learning is to run a query with a filter and a sensible limit and count the rows returned. Recent writes may take a moment to be visible, so for a precise number right after inserting, create the collection with Strong consistency.
Partitions, briefly. A collection can be divided into partitions, which are named subsets, and you can search only some of them. A new collection always has one called _default. Beginners rarely need more, and Milvus Lite does not support them at all. When you eventually need to separate data by customer or by tenant, a feature called the partition key does this more cleanly, and the Mid-level part covers it.
exp_minilm_cosine. Collections are cheap. Dropping a whole collection and rebuilding it is far less error-prone than trying to repair one in place, and it mirrors how you will later re-index when you change embedding models.
- List your collections and describe one of them.
- Release the collection, then search it and read the error.
- Load it again and search successfully.
Putting it all together: a small semantic search project
Time to build something complete. We will make a tiny question-and-answer search over a handful of sentences, using only what you have learned. To keep the code runnable without downloading a model, we use a deliberately simple stand-in for an embedding model: it hashes each word into a slot of a 64-number vector and normalises the result. This is a toy. It matches shared words rather than meaning, so it will not find synonyms. Its job is to let you run the whole pipeline now. In a real project you replace one function with a call to a genuine embedding model and change nothing else except the dimension.
Step one: the embedding function and the data
import hashlib
import math
DIM = 64
def embed(text: str) -> list[float]:
"""Toy embedding: hash each word into one of DIM slots, then normalise."""
vec = [0.0] * DIM
for word in text.lower().split():
word = word.strip(".,?!")
slot = int(hashlib.md5(word.encode("utf-8")).hexdigest(), 16) % DIM
vec[slot] += 1.0
norm = math.sqrt(sum(x * x for x in vec)) or 1.0
return [x / norm for x in vec]
docs = [
{"id": 1, "text": "Reset your password from the account settings page.", "section": "account"},
{"id": 2, "text": "Invoices are emailed on the first day of each month.", "section": "billing"},
{"id": 3, "text": "You can change your billing address in the account settings page.", "section": "billing"},
{"id": 4, "text": "Our support team answers email within one business day.", "section": "support"},
{"id": 5, "text": "Delete your account from the account settings page.", "section": "account"},
]
Step two: build the collection
from pymilvus import MilvusClient, DataType
client = MilvusClient("./support_demo.db") # or uri="http://localhost:19530"
if client.has_collection("support_docs"):
client.drop_collection("support_docs")
schema = MilvusClient.create_schema(auto_id=False, enable_dynamic_field=False)
schema.add_field("id", DataType.INT64, is_primary=True)
schema.add_field("vector", DataType.FLOAT_VECTOR, dim=DIM)
schema.add_field("text", DataType.VARCHAR, max_length=2000)
schema.add_field("section", DataType.VARCHAR, max_length=64)
index_params = client.prepare_index_params()
index_params.add_index(field_name="vector", index_type="AUTOINDEX", metric_type="COSINE")
client.create_collection(
collection_name="support_docs",
schema=schema,
index_params=index_params,
)
We turned the dynamic field off so that a typo in a field name is an error rather than a silently stored extra. Because we passed index_params, the collection is indexed and loaded when the call returns.
Step three: insert and search
rows = [{**d, "vector": embed(d["text"])} for d in docs]
client.insert(collection_name="support_docs", data=rows)
def ask(question: str, section: str | None = None, k: int = 2):
kwargs = {}
if section:
kwargs["filter"] = "section == {s}"
kwargs["filter_params"] = {"s": section}
hits = client.search(
collection_name="support_docs",
data=[embed(question)],
limit=k,
output_fields=["text", "section"],
**kwargs,
)
for hit in hits[0]:
print(f'{hit["distance"]:.2f} [{hit["entity"]["section"]}] {hit["entity"]["text"]}')
ask("how do I change my account settings")
print("--- billing only ---")
ask("how do I change my account settings", section="billing")
You should see account and billing sentences about the settings page ranked first, because they share the words "account", "settings" and "page" with the question. The second call restricts the search to the billing section, so only billing rows can appear however well another section matches. If a hit's text is not what you expect, remember what the toy embedding does: it counts shared words. With a real embedding model the first query would also find "Reset your password" through meaning alone.
Step four: maintain the data
client.upsert(collection_name="support_docs", data=[{
"id": 4,
"vector": embed("Our support team answers email within four business hours."),
"text": "Our support team answers email within four business hours.",
"section": "support",
}])
client.delete(collection_name="support_docs", filter="section == 'billing'")
print(client.query(collection_name="support_docs", filter="id >= 0",
output_fields=["id", "section"], limit=10))
The upsert replaces row 4 by its primary key. The delete removes both billing rows. The final query lists what remains, which should be ids 1, 4 and 5.
Step five: swap in a real model
The only function that changes is embed. Choose an embedding model, set DIM to the length of its output, and create a new collection, since the dimension of an existing one cannot change. Keep the model name written down next to the collection. After that swap the search finds sentences by meaning, and the toy's limitations disappear. The retrieval step of a RAG system is exactly this code: embed the user's question, search, and pass the returned texts to a language model. Tools such as LangChain and LlamaIndex wrap these calls, and it helps to have written them once by hand.
- Run all five steps against Milvus Lite or Docker standalone.
- Ask a question that shares no words with any document and see how the toy embedding fails.
- Write down which single function you would replace to fix it, and what else must change.
embed function, plus the DIM constant and a fresh collection. The database code stays identical, which is the point: Milvus is independent of how you produce vectors.
What you can now do, and what comes next
You can start Milvus in the mode that fits your situation, connect with MilvusClient, and explain the five nouns: collection, field, entity, index, and metric. You can create a collection in one call or define a custom schema with an index, and you know that loading is what makes it searchable. You can insert in batches, upsert safely, delete by id or filter, search with and without filters, query by scalar conditions, and read the most common errors without panic. You also know the two decisions that matter most before you scale: the embedding model fixes the dimension, and the defaults leave the server unauthenticated.
Where you will want to go next, in the Mid-level part: choosing and tuning an index such as HNSW or IVF rather than leaving it to AUTOINDEX, hybrid search that combines vector similarity with full-text keyword matching, partition keys for multi-tenant data, consistency levels in depth, schema changes on live collections, and server-side embedding functions. After that, the Senior part covers architecture, failure modes, scaling, security, upgrades, and operating Milvus as a platform. Neighbouring guides worth reading alongside: Qdrant, pgvector, and Chroma show how other vector stores answer the same questions, which is the best way to decide whether Milvus is the right fit; Ragas shows how to measure whether retrieval is actually good.
Sources
- Milvus documentation home: https://milvus.io/docs
- Quickstart: https://milvus.io/docs/quickstart.md
- Milvus Lite: https://milvus.io/docs/milvus_lite.md
- Prerequisites: https://milvus.io/docs/prerequisite-docker.md
- Install standalone with Docker: https://milvus.io/docs/install_standalone-docker.md
- Install standalone with Docker Compose: https://milvus.io/docs/install_standalone-docker-compose.md
- Install standalone on Windows: https://milvus.io/docs/install_standalone-windows.md
- Create a collection: https://milvus.io/docs/create-collection.md
- Filtering expressions: https://milvus.io/docs/boolean.md
- Filter templating: https://milvus.io/docs/filtering-templating.md
- Consistency levels: https://milvus.io/docs/consistency.md
- Index explained: https://milvus.io/docs/index-explained.md
- Authenticate user access: https://milvus.io/docs/authenticate.md
- Limits: https://milvus.io/docs/limitations.md
- Troubleshooting: https://milvus.io/docs/troubleshooting.md
- Release notes: https://milvus.io/docs/release_notes.md
- Attu: https://github.com/zilliztech/attu