تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
ChromaLLMsVector databases3 مستويات102 قسمًايغطّي Chroma 1.5دليل بالإنجليزية

The Complete Chroma Guide

Add embeddings search to your app in minutes with the open-source Chroma database. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
17sections
22examples

This is part one of three. It covers everything you need to do real work with Chroma, not a teaser. By the end you can store text in a collection, search it by meaning, narrow the search with filters, keep your data on disk between runs, move from a script to a running server, and read the error messages you will meet along the way. Mid-level and Senior take the same topics further; nothing here is thrown away.

Every command and example was checked against Chroma 1.5.9, the current stable release. Pin that version when you follow along. Each section ends with a Try it task. Do them as you go: search by meaning only feels real once you have watched your own sentences come back in a surprising but sensible order.

What Chroma is, and the problem it solves

Chroma is an open-source database built for retrieval: you give it pieces of text (or other data), and later you ask a question and it hands back the pieces most relevant to that question. It is licensed Apache-2.0, and its own documentation describes it as open-source data infrastructure for AI.

To see why this needs a special kind of database, start with how ordinary search works. A classic database finds rows where a column equals a value, or where a string contains a word. If you search a help centre for "my card was declined" and the relevant article says "payment failed", a keyword search finds nothing, because the two sentences share no words. A human sees at once that they mean the same thing. Chroma exists to close that gap.

The trick is the embedding. An embedding is a list of numbers, a vector, that a machine learning model produces from a piece of text so that texts with similar meaning get similar numbers. "My card was declined" and "payment failed" end up close together in that number space, even though they share no words. Searching then becomes geometry: turn the question into a vector, find the stored vectors nearest to it, and return the text they came from.

TEXTyour documents
→
EMBEDDINGa list of numbers
→
CHROMAstores and indexes
→
QUERYnearest neighbours

The diagram is the whole idea. Everything else in this guide is detail about one of those four boxes.

Why does anyone need this now? Large language models (LLMs) answer from what they learned in training, which does not include your company's contracts, your wiki, or yesterday's tickets. The standard remedy is retrieval-augmented generation, usually shortened to RAG: before asking the model a question, you first retrieve the few passages from your own data that are relevant, and paste them into the prompt. Chroma is the retrieval half. It is not the model, and it does not write answers. It finds the right passages, quickly, so the model has something true to work from.

What people use it for:

📚

Question answering over documents

Index manuals, policies or papers, then retrieve the passages that answer a question before calling an LLM.

🔎

Semantic search

Find tickets, products or messages by what they mean, not by the exact words used.

🧠

Agent memory

Let an assistant store what it learned and recall it later in a different conversation.

🧪

Prototyping

Run it inside your own script with no server and no account, then move to a server when the project grows.

Chroma runs in three ways, and it helps to know all three from the start. Local means it runs inside your Python process as a library, which is ideal for learning and prototypes. Single-node means one Chroma server that your applications talk to over the network; the documentation positions this for small to medium workloads, typically fewer than ten million records across a handful of collections. Distributed means a multi-service deployment for very large workloads, and Chroma Cloud is the managed version of that. This guide teaches local and single-node, which is where every learner should begin. The API you learn is almost the same in all of them.

Try it
  1. Write three short sentences that mean the same thing as each other but share no important words, for example about a refund, a late delivery, or a forgotten password.
  2. Write one unrelated sentence, such as a recipe step.
  3. Keep these four sentences. You will search them in the first project.
a tiny test set where a human can see which sentences belong together. If Chroma ranks them the same way you would, search by meaning is working.

The model: tenants, databases, collections and records

Chroma has a small number of nouns. Learn them now and the API will read like plain English.

A record is one stored item. It has a required id, which is a string you choose and which must be unique within its collection. It usually has a document, the text itself. It has an embedding, the vector for that text. It may have metadata, a flat dictionary of small facts about the item such as its source file, page number or author. It can also carry a uri, a pointer to where the original lives. Beginners work mostly with id, document and metadata, and let Chroma compute the embedding.

A collection is a named group of records. It is the fundamental unit of storage and querying in Chroma, and each collection is indexed independently. If you are coming from relational databases, a collection is roughly a table, though a better mental picture is a folder of related notes that you search together. You would keep your support articles in one collection and your product descriptions in another, because you rarely want a question about a refund to return a product description.

Above the collection sit two larger containers. A database is a logical namespace for an application or environment, for instance one for testing and one for production. A tenant is the top-level unit, standing for a user, team or account. When you create a client without specifying either, Chroma uses default_tenant and default_database, and you can ignore both for a long time. They matter later, when one server hosts several teams, and the senior track covers them.

TENANTan account or team
→
DATABASEan app or environment
→
COLLECTIONrecords searched together
→
RECORDid, document, metadata

Two rules about collections catch beginners early. First, a collection name must be between 3 and 512 characters, drawn from letters, digits, dots, underscores and hyphens, and it must start and end with a letter or digit. A name like a is rejected, as is -notes. Second, every embedding in a collection must have the same number of dimensions, and the first record you add fixes that number. This is why you cannot mix vectors from two different embedding models in one collection, a point we return to in the errors section.

There is one more noun worth meeting now: the client. A client is the object your code uses to talk to Chroma. You create one, and then every operation, such as creating a collection, goes through it. Which client you create decides where the data lives, and that is the next decision after you install the tool.

Try it
  1. Pick a project idea, such as a recipe finder or a company FAQ.
  2. Decide what a single record is, and what its id, document and one or two metadata fields would be.
  3. Decide whether you need one collection or several.
a clear sentence like "one record is one FAQ answer, the id is its slug, the metadata is its topic". Good records are small enough to be one idea each.

Embeddings, distance and the embedding function

You can use Chroma for a long time without understanding embeddings in depth, but three ideas will save you from confusing bugs.

The first is the embedding function, shortened to EF. An embedding function is the piece of code that turns text into a vector. In Chroma you attach one to a collection, and then add and query call it for you whenever you hand over plain text. If you never specify one, Chroma uses its default, a small, fast, well-known model called all-MiniLM-L6-v2 from the Sentence Transformers family, which produces vectors of 384 numbers. In Python it runs locally through ONNX, so no API key is needed. The first time it is used, it downloads the model to ~/.cache/chroma/onnx_models. That one-time download is why your very first add takes longer than the rest, and why it fails on a machine with no internet access.

You can replace the default with a model from a provider such as OpenAI, Cohere, Google, Hugging Face, Ollama, Mistral, Jina or Voyage. Chroma ships ready-made embedding functions for these, and you pass one when you create the collection. You can also skip embedding functions entirely and supply your own vectors, in which case Chroma stores them as given. The beginner path is to use the default and revisit the choice once the first project works.

The second idea is distance. When you query, Chroma returns records nearest to your question, and "nearest" needs a definition. Chroma offers three: l2 (squared Euclidean distance), cosine (one minus cosine similarity) and ip (one minus the inner product). The default is l2. For all of them, a lower number means more similar, so the best match has the smallest distance. Many learners expect a high score to mean a good match, as in a percentage, and misread the output by exactly that inversion. For text embeddings, cosine is the most common choice by convention, and you choose it when you create the collection. The space is fixed at creation and cannot be changed later, which is worth remembering before you load a million records.

The third idea is that the same model must embed both the stored text and the question. A vector is only meaningful relative to the model that produced it. If you add documents with one model and query with another, you will either hit a dimension error or, worse, get plausible-looking nonsense when the dimensions happen to match. Chroma helps by remembering the embedding function inside the collection configuration since version 1.1, so a collection you reopen later uses the same one automatically. API keys are not stored, though. If your embedding function needs one, such as OPENAI_API_KEY, every process that uses the collection must have that environment variable set.

Lower distance is better A query result with distances [0.31, 0.88, 1.40] means the first document is the closest match. Distances are not percentages and are not capped at 1. Compare them with each other within one query rather than reading them as an absolute score.
Try it
  1. With pencil and paper, rank your four test sentences from the previous section for the question "I want my money back".
  2. Write down your ranking, best match first.
  3. Keep it, and compare it with Chroma's ranking in the first project.
your own judgement as a baseline. The goal is not that Chroma matches you exactly, but that the obviously relevant sentences land near the top.

Installing Chroma and checking the setup

Chroma installs as an ordinary Python package. Use a virtual environment so the install stays tidy, and pin the version so your results match this guide.

BASH
python -m venv .venv
source .venv/bin/activate          # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install "chromadb==1.5.9"

The package requires Python 3.9 or newer, and Python 3.14 has been supported since 1.5.3. Prebuilt wheels include the compiled Rust core, so you do not normally need a compiler. Installing chromadb also installs the chroma command-line tool, which you will use to run a server and to browse your data. If your system blocks global pip installs, pipx install chromadb is the alternative for the command-line tool.

Verify the setup with two commands:

BASH
chroma --version
python -c "import chromadb; print(chromadb.__version__)"

The second command should print 1.5.9. If it prints a different number, you may have an older install shadowing the new one, usually because you forgot to activate the virtual environment.

There are other ways in, and you will meet them in tutorials. JavaScript and TypeScript developers use npm install chromadb @chroma-core/default-embed, and the JavaScript client always needs a running server. There is a standalone command-line installer for people without Python or Node, and an official Docker image, chromadb/chroma:1.5.9. There is also a thin Python package, chromadb-client, meant for applications that only talk to a remote server; install it instead of chromadb, not alongside it, and remember that it has no default embedding function. This guide uses the full chromadb package throughout.

Three platform notes will save you an afternoon. Chroma needs SQLite newer than 3.35; on an old Linux system the official troubleshooting page suggests pip install pysqlite3-binary, and inside Docker you should use a Debian bookworm base or newer. If a build ever fails with Failed to build hnswlib, you are building from source rather than using a wheel, and setting HNSWLIB_NO_NATIVE=1 or installing the Xcode command-line tools on macOS usually helps. And if your employer routes traffic through a TLS-inspecting proxy, the first model download can fail with [SSL: CERTIFICATE_VERIFY_FAILED] certificate verify failed: self-signed certificate in certificate chain. The fix is to make Python trust your company's root certificate, for example through the SSL_CERT_FILE environment variable, or to copy the model into ~/.cache/chroma/onnx_models from a machine that can download it. Do not turn certificate verification off.

Heartbeat is not the package version If you later run a server and ask it for its version over HTTP, a 1.5.9 server answers "1.0.0". That number is the server API version, not the package you installed. It does not mean you installed the wrong release. The package version is the one printed by chromadb.__version__.
Try it
  1. Create a folder called chroma-lab, make a virtual environment inside it, and activate it.
  2. Install chromadb==1.5.9.
  3. Run both verification commands.
1.5.9 printed by both. If chroma --version is not found, your virtual environment is not active.

Your first project: a searchable knowledge base

Now build something. We will store a handful of help-centre sentences, search them by meaning, and read the results. Create a file called first.py.

first.py
import chromadb

client = chromadb.EphemeralClient()
collection = client.create_collection(name="help_centre")

collection.add(
    ids=["a1", "a2", "a3", "a4"],
    documents=[
        "Your payment failed because the card was declined.",
        "To reset your password, use the link on the sign-in page.",
        "Refunds are issued to the original payment method within five days.",
        "Preheat the oven to 180 degrees before baking the bread.",
    ],
    metadatas=[
        {"topic": "billing"},
        {"topic": "account"},
        {"topic": "billing"},
        {"topic": "recipes"},
    ],
)

results = collection.query(
    query_texts=["I want my money back"],
    n_results=2,
)
print(results)

Run it with python first.py. The first run pauses while the default embedding model downloads, then prints a dictionary. Let us walk through what happened, because each line teaches one idea.

chromadb.EphemeralClient() creates a client that keeps everything in memory. When the script ends, the data is gone. That makes it the safest place to experiment, because there is nothing to clean up. create_collection(name="help_centre") creates a collection with the default embedding function and the default distance.

collection.add(...) stores four records. We passed ids, documents and metadatas as three parallel lists, so the first id goes with the first document and the first metadata dictionary. We never passed embeddings. Because the collection has an embedding function and we gave it text, Chroma computed the vectors for us. This is the moment the model download happens.

collection.query(query_texts=[...], n_results=2) embeds the question with the same model and returns the two nearest records. Notice that query_texts is a list, even for one question. Chroma can answer several questions in one call, and that shape explains the odd nesting of the output.

The result looks roughly like this (your distances will differ slightly by platform):

PYTHON
{
  'ids': [['a3', 'a1']],
  'documents': [['Refunds are issued to the original payment method within five days.',
                 'Your payment failed because the card was declined.']],
  'metadatas': [[{'topic': 'billing'}, {'topic': 'billing'}]],
  'distances': [[0.95, 1.21]],
  'embeddings': None, 'uris': None, 'data': None, 'included': [...]
}

Read it like this. Each key holds a list, and each of those holds one inner list per question you asked. We asked one question, so every outer list has one element, which is itself a list of two matches, best first. That is why you write results["documents"][0] to get the documents for your first question. The refund sentence ranks first, though it shares no important word with "I want my money back", which is the whole point of semantic search. The recipe about baking bread did not appear at all. Embeddings show as None because they are left out by default to keep responses small; you can ask for them explicitly.

Print one field at a time When you are learning, print results["documents"][0] and results["distances"][0] separately instead of the whole dictionary. The nested lists stop looking scary once you realise that [0] simply means "the answers for my first question".
Try it
  1. Replace the four documents with your own test sentences from the first section.
  2. Run three different questions, one at a time, and print only the documents.
  3. Change n_results to 4 and watch the unrelated sentence arrive last.
related sentences first and the unrelated recipe last. If your own ranking and Chroma's disagree, that is normal for a small default model; note which cases surprise you.

Choosing a client: memory, disk or server

The client decides where your data lives, so choosing it is the first real design decision. Chroma offers four you will use as a beginner, and they all return the same kind of object with the same methods.

EphemeralClient() keeps data in memory. chromadb.Client() is equivalent. Use it for experiments, tests and notebooks where you want a clean slate each run.

PersistentClient(path="./chroma") stores data on disk in the folder you name. Run your script twice and the second run sees the first run's data. Write path explicitly. The official documentation text mentions a .chroma default, but in 1.5.9 the actual default is ./chroma, and relying on a default you cannot see is how people end up with several copies of a database in different working directories.

HttpClient(host="localhost", port=8000) connects to a Chroma server running elsewhere, which might be a second terminal on your own machine or another machine entirely. There is an AsyncHttpClient with the same arguments for asynchronous code. CloudClient() connects to Chroma Cloud, reading CHROMA_API_KEY, CHROMA_TENANT and CHROMA_DATABASE from the environment, but you do not need an account to learn Chroma, and it is not covered in this beginner guide beyond this mention.

Here is the same idea as a persistent script, which is the version most real projects start with:

persistent.py
import chromadb

client = chromadb.PersistentClient(path="./chroma")
collection = client.get_or_create_collection(name="help_centre")

print("records stored:", collection.count())

collection.upsert(
    ids=["a1"],
    documents=["Your payment failed because the card was declined."],
    metadatas=[{"topic": "billing"}],
)
print("records now:", collection.count())

Run it twice. The first run prints records stored: 0 then records now: 1. The second prints 1 and 1, because the data survived and upsert replaced the existing record instead of duplicating it. After running, look in your working directory: a chroma folder appeared, holding a SQLite file called chroma.sqlite3 plus index directories. That folder is your database. To back it up, stop your script and copy the whole folder.

Persistent client

  • Data survives between runs
  • One folder holds everything
  • No server to start or secure
  • Use for apps, notebooks you keep, small tools

Ephemeral client

  • Data vanishes when the process ends
  • Nothing to clean up
  • Perfect for tests and quick experiments
  • Never use it for data you want to keep
Do not open one folder from two processes A persistent client is meant to be used by a single process at a time. If two scripts open the same ./chroma folder concurrently you risk corruption or confusing behaviour. When several programs need to share data, run a Chroma server and have each connect with HttpClient.
Try it
  1. Run persistent.py twice and confirm the second run reports one stored record.
  2. Delete the chroma folder and run it again.
  3. Change the client line to EphemeralClient() and notice that the count is always zero at the start.
proof that the folder is the database. Deleting it resets everything, which is a useful and a dangerous fact.

Working with collections

Collections are the objects you create, find, list and delete. The calls are few, and each has one behaviour you should know.

PYTHON
col = client.create_collection(name="notes")          # fails if it exists
col = client.get_collection(name="notes")             # fails if it does not exist
col = client.get_or_create_collection(name="notes")   # the safe default for scripts
client.list_collections()                             # a list of Collection objects
client.count_collections()                            # how many exist
client.delete_collection(name="notes")                # irreversible
col.count()                                           # records in this collection
col.peek(limit=5)                                     # look at the first few records

create_collection raises an error if the name is taken: Collection [notes] already exists. get_collection raises Collection [notes] does not exist if it is missing. For scripts you run repeatedly, get_or_create_collection is the forgiving choice, but know its catch: if the collection already exists, any other arguments you pass, such as a new distance setting or new metadata, are ignored. It will not overwrite what is there.

You set the distance and a short description when you create a collection:

PYTHON
col = client.create_collection(
    name="notes",
    metadata={"description": "Course notes"},
    configuration={"hnsw": {"space": "cosine"}},
)

metadata here is a free-form note about the collection. configuration carries index settings, and {"hnsw": {"space": "cosine"}} chooses cosine distance. Remember that the space is fixed at creation. If you picked the wrong one, the fix is to create a new collection and add the data again, not to modify the old one.

list_collections() returns Collection objects, which is easy to forget because some old tutorials describe a version that returned only names. In 1.x you get objects, so print [c.name for c in client.list_collections()] when you only want names. It returns up to 100 by default, and you can pass limit and offset to page through more. peek is a quick way to see what is inside without writing a query, and count is how you check that a load worked.

You can rename a collection or change its metadata with col.modify(name="new-name", metadata={...}). And delete_collection is permanent, with no confirmation and no undo, so treat it as you would rm -r. If you are scripting cleanup, delete by exact name and print what you are about to delete first.

One more thing: when you attach your own embedding function, or choose to bring your own vectors, you say so at creation. Passing embedding_function=None creates a collection with no embedding function, which means you must supply embeddings yourself on every add and query. This is a legitimate pattern when you compute vectors elsewhere, and it explains one error you will see in the errors section.

Try it
  1. Create a collection called notes with cosine distance.
  2. Call create_collection with the same name again and read the error.
  3. Call get_or_create_collection instead, then list all collection names and delete notes.
an already exists error on the second create, then a clean run with the forgiving call.

Adding, updating and deleting records

Reading is the exciting part, but most bugs live in writing, because Chroma is forgiving in ways that hide mistakes. Learn the four write operations and their quirks.

PYTHON
col.add(ids=["n1", "n2"], documents=["first note", "second note"],
        metadatas=[{"week": 1}, {"week": 2}])

col.update(ids=["n1"], documents=["first note, edited"])

col.upsert(ids=["n2", "n3"], documents=["second note", "third note"],
           metadatas=[{"week": 2}, {"week": 3}])

col.delete(ids=["n3"])

add inserts new records. You must provide documents, embeddings, or both; metadata is optional. The ids must be strings, and they must be unique inside the call. The quirk that surprises everyone: adding an id that already exists is silently ignored. No error, no overwrite, and your new text is simply not stored. If you re-run a loading script and your edits never seem to take effect, this is why.

update changes records that already exist. If you pass new documents without embeddings, Chroma re-embeds them. Updating an id that does not exist does not raise; an error is logged and nothing changes, so check your ids if an update seems to vanish.

upsert means update-or-insert, and it is the right default for any script you will run more than once, because running it twice leaves the same result as running it once. That property has a name, idempotent, and it is worth wanting in every loading job you ever write.

delete removes records by id, or by a metadata filter such as col.delete(where={"week": 3}). Deleting by filter is powerful and easy to get wrong, so run the same filter through col.get(where=...) first to see what would go.

The rules about what you can store are strict, and breaking them produces errors you will meet in the next sections. Ids must be strings: 1 is rejected, "1" is fine. Metadata must be flat: values may be strings, integers, floats or booleans, or non-empty lists of those, but never a nested dictionary and never None in a call to add. If your data is nested, either flatten the keys, writing author_name rather than {"author": {"name": ...}}, or serialise the nested part to a JSON string. Lists of parallel inputs must be the same length, so three ids with two documents is an error.

Chroma also limits how many records you can send at once. Ask the client with client.get_max_batch_size(), which returns 5461 on a local 1.5.9 install; sending more raises Batch size 5462 exceeds maximum batch size 5461. For speed on larger loads, the documentation suggests batches in the range of 50 to 250 and sending several in parallel. For now, a simple loop that writes a few hundred records at a time is plenty.

Use stable, meaningful ids An id like faq-refunds-001 or a hash of the source text lets you re-run a load safely with upsert and lets you find a record again. Random ids make it impossible to tell whether a record is already stored, and you end up with duplicates.
Try it
  1. Add a record with id x1 and the document "original".
  2. Add id x1 again with the document "changed" and then get it. Which text is stored?
  3. Now upsert the same id with "changed" and get it again.
"original" after the second add, because the duplicate was silently ignored, and "changed" only after the upsert. This single experiment prevents a whole class of confusion.

Querying and getting: two ways to read

Chroma has two read operations, and picking the right one is a beginner skill. query ranks records by similarity to a question. get retrieves records by id or by a filter, with no ranking at all. Use query when you want what is closest in meaning, and get when you know exactly which records you want, or want to list everything that matches a condition.

PYTHON
# Ranked by meaning
res = col.query(
    query_texts=["how do I get a refund"],
    n_results=3,
    include=["documents", "metadatas", "distances"],
)

# Exact lookup, no ranking
one = col.get(ids=["n1"])
page = col.get(limit=10, offset=0, include=["documents"])

The n_results argument defaults to 10. If you ask for more results than the collection has, 1.5.9 simply returns everything it has with no error, but asking for zero is an error: Number of requested results 0, cannot be negative, or zero.

The include argument controls which fields come back. For query the default is documents, metadatas and distances. For get it is documents and metadatas. Valid values are documents, embeddings, metadatas, distances, uris and data. Ids are always returned and must not be listed. Writing include=["ids", ...] is a common mistake and produces Expected include item to be one of documents, embeddings, metadatas, distances, uris, data, got ids. Embeddings are left out by default because they are large; ask for them only when you need them, and in Python they arrive as NumPy arrays.

The two operations differ in the shape of their output, and mixing them up causes most beginner indexing errors. query handles a list of questions, so its results are nested: results["ids"][0] holds the ids for the first question. get has no questions, so its results are flat: one["ids"] is a plain list. Results are also column-major: instead of a list of record objects, you get one list of ids, one list of documents and so on, all in the same order. To walk through matches as rows, zip the columns together:

PYTHON
res = col.query(query_texts=["refund"], n_results=3)
for doc_id, doc, dist in zip(res["ids"][0], res["documents"][0], res["distances"][0]):
    print(f"{dist:.3f}  {doc_id}  {doc}")

One honest warning about search quality. Chroma returns the nearest records even when none of them is actually relevant. If you ask about quantum physics in a cooking collection, you still get results, just with large distances. A search with no good answer never comes back empty. In a real application you therefore look at the distances, and decide on a cut-off that suits your data and your model. Beginners often assume a result means a match; it only means "the least bad".

Try it
  1. Run a query that has an obvious answer in your collection, and print the distances.
  2. Run a query about something completely unrelated and print the distances again.
  3. Compare the two sets of numbers and note roughly where relevant stops and irrelevant starts.
clearly larger distances for the unrelated question, though results still come back. That gap is how you would build a sensible threshold later.

Filtering with where and where_document

Pure similarity is rarely enough. You often want "the closest passages, but only from the billing section" or "only from this year". Chroma lets you combine similarity with filters, and the filter runs together with the search rather than after it.

There are two filters. where filters on metadata. where_document filters on the text of the document itself. Both work on query, get and delete.

The shortest where is an equality test:

PYTHON
col.query(query_texts=["card problem"], n_results=3, where={"topic": "billing"})
col.get(where={"week": 2})

Beyond equality, where supports comparison and set operators, each written as a dictionary key beginning with a dollar sign:

PYTHON
col.get(where={"week": {"$gte": 2}})                       # greater than or equal
col.get(where={"topic": {"$in": ["billing", "account"]}})  # one of several values
col.get(where={"$and": [{"topic": "billing"}, {"week": {"$lt": 3}}]})
col.get(where={"$or": [{"topic": "billing"}, {"topic": "account"}]})

The full list is $eq, $ne, $gt, $gte, $lt, $lte, $in and $nin, plus $and and $or for combining conditions. There are also $contains and $not_contains for metadata values that are lists. A bare {"field": value} is shorthand for $eq. Two traps: an empty filter, where={}, is rejected with Expected where to have exactly one operator, got {}, so leave where out instead of passing an empty dictionary; and a typo such as $foo produces a clear error listing the valid operators.

where_document searches inside the text:

PYTHON
col.query(query_texts=["refund"], n_results=3,
          where_document={"$contains": "five days"})
col.get(where_document={"$regex": "^Refunds"})

It supports $contains, $not_contains, $regex and $not_regex, combined with $and and $or. Text matching here is case-sensitive, so "Refund" and "refund" are different. Use this filter for exact words, codes or phrases that a semantic search might blur; it is the beginner's version of combining keyword and meaning search.

How does a filter interact with ranking? Chroma applies the filter and ranks only within the records that pass. This means a very restrictive filter can leave fewer records than n_results, and then you simply get fewer results back. In rare, heavily filtered cases against a large index you may see the HNSW message Cannot return the results in a contiguous 2D array. Probably ef or M is too small; the first remedy is to ask for fewer results.

Decide your metadata before you load A filter can only use metadata you stored. If you think you will ever want to search by source, date, author or language, put it in the metadata when you add the record. Adding it later means updating every record.
Try it
  1. Add six records with a topic and a numeric week in the metadata.
  2. Query with a where on topic only, then add a $gte condition on week using $and.
  3. Try where={} on purpose and read the error.
shrinking result sets as you add conditions, and a clear error for the empty filter.

Running Chroma as a server and browsing your data

A persistent client is perfect until a second program needs the data. Then you want a server. Chroma's server is one command:

BASH
chroma run --path ./chroma-data --port 8000

The startup banner prints Saving data to: ./chroma-data and Connect to Chroma at: http://localhost:8000. The chroma run command binds to localhost by default, so only your own machine can reach it; add --host 0.0.0.0 only when you truly intend other machines to connect, because, as the errors and security notes say, a 1.x server has no built-in authentication. The signature is chroma run [CONFIG_PATH] [--path <dir>] [--host localhost] [--port 8000], and you can supply a YAML configuration file instead of flags.

With the server running, check it from a second terminal:

BASH
curl http://localhost:8000/api/v2/heartbeat

You get back a JSON object with a nanosecond timestamp, which proves the server is alive. Note the /api/v2/ in the path. The older /api/v1/ endpoints were removed in 1.0 and now return HTTP 410 Gone with the message The v1 API is deprecated. Please use /v2 apis. If you ever see that, an old client or an old tutorial is talking to a new server.

Connect from Python with the HTTP client. Your add and query code is unchanged:

server_client.py
import chromadb

client = chromadb.HttpClient(host="localhost", port=8000)
print(client.heartbeat())            # an integer, nanoseconds

col = client.get_or_create_collection(name="help_centre")
col.upsert(ids=["a1"], documents=["Your payment failed."], metadatas=[{"topic": "billing"}])
print(col.query(query_texts=["card problem"], n_results=1)["documents"][0])

If the server is not running you get ValueError: Could not connect to a Chroma server. Are you sure it is running?. The first checks are always the same: is the server process alive, is the port right, and, with Docker, did you publish the port?

Docker is the most common way to run a server you want to leave running:

BASH
docker run -v ./chroma-data:/data -p 8000:8000 chromadb/chroma:1.5.9

The -v flag mounts a folder at /data, which is where the container stores its database, and -p 8000:8000 publishes the port. Since 1.0, the data path inside the container is /data; older guides that mount /chroma/chroma are out of date, and following them silently stores your data inside the disposable container. If you are new to containers, the Docker guide explains images, volumes and ports from scratch.

Finally, you can look at your data in a browser interface. The command-line tool includes chroma browse:

BASH
chroma browse help_centre --local
chroma browse help_centre --path ./chroma

Pointing it at a collection lets you scroll records and see metadata, which beats printing dictionaries when you are checking whether a load worked. The same tool offers chroma install --list for sample apps, chroma update to upgrade the command-line tool, and chroma docs to open the documentation.

No built-in login on a self-hosted server Chroma 1.x removed its built-in authentication. Anyone who can reach the port can read, change and delete everything, including whole collections. Keep a self-hosted server on a private network or behind a proxy that enforces authentication, and never expose port 8000 to the internet as is. This matters for Gulf and Egyptian employers with data-residency rules: where the server runs is where your data lives.
Try it
  1. Start chroma run --path ./chroma-data --port 8000 in one terminal.
  2. Run curl against the heartbeat endpoint from another terminal.
  3. Run server_client.py, stop the server, and run it again to read the connection error.
a working heartbeat, a working query, then the Could not connect error once the server is gone.

Bringing your own embeddings and choosing an embedding function

So far Chroma has computed every vector with its default model. Two other patterns come up quickly, and both are one argument away.

The first is choosing a different embedding function. Chroma ships wrappers in chromadb.utils.embedding_functions for many providers. For example, with OpenAI:

PYTHON
from chromadb.utils import embedding_functions

ef = embedding_functions.OpenAIEmbeddingFunction(model_name="text-embedding-3-small")
col = client.create_collection(name="docs_openai", embedding_function=ef)

This reads your key from the OPENAI_API_KEY environment variable; if your key lives under another name, pass api_key_env_var="MY_VAR". Never paste a key into source code. There are wrappers for Cohere, Google, Hugging Face, Ollama, Mistral, Jina, Voyage and others, and a local option, SentenceTransformerEmbeddingFunction, that runs a model on your own machine. Pick a provider based on cost, language coverage and privacy. For Arabic content in particular, test a multilingual model against your own sentences before committing, because a small English-centred model may rank Arabic text poorly. The default model is mainly trained on English, and you should verify it on your data rather than assume.

The second pattern is supplying embeddings yourself, which you do when you compute vectors in another system:

PYTHON
col = client.create_collection(name="external", embedding_function=None)
col.add(
    ids=["e1", "e2"],
    embeddings=[[0.1, 0.2, 0.3], [0.2, 0.1, 0.4]],
    documents=["first", "second"],
)
print(col.query(query_embeddings=[[0.1, 0.2, 0.35]], n_results=1)["ids"])

Chroma stores the vectors exactly as given and never re-embeds them. Note the dimension: the first add fixed it at three, so a later query with a two-number vector raises Collection expecting embedding with dimension of 3, got 2. And on a collection created with embedding_function=None, calling query(query_texts=...) fails with You must provide an embedding function to compute embeddings, because there is nothing to turn the text into numbers. Either pass query_embeddings or attach a function.

Once a collection has an embedding function, Chroma remembers its configuration, so get_collection later rebuilds it without you passing it again, provided the client and server are at least version 1.1. What it cannot remember is your secret key. When you move a script to a new machine, set the environment variable there too. If you change the embedding model for a collection that already has data, the old vectors no longer match the new ones. The only correct fix is a new collection and a full reload.

Try it
  1. Create a collection with embedding_function=None and add three records with hand-made three-number vectors.
  2. Query it with a three-number vector, then with a two-number vector.
  3. Query it with query_texts and read the error.
one clean result and two clear errors, which teaches how tightly vectors, dimensions and embedding functions are tied together.

Configuration, settings and housekeeping

Most beginners never need to touch configuration, but you should know what exists so that you can recognise it when it appears in other people's code.

Per-collection configuration controls the index. The common keys sit under hnsw: space (the distance), ef_construction, ef_search, max_neighbors, resize_factor, sync_threshold, num_threads and batch_size. The 1.5.9 defaults include space of l2, ef_construction of 100, ef_search of 100 and max_neighbors of 16. As a beginner, set only space, and leave the rest alone. Two are creation-only: the space and ef_construction cannot change afterwards. Trying to change the space with modify produces unknown field space, expected one of ef_search, max_neighbors, num_threads, resize_factor, sync_threshold, batch_size. The remedy is a new collection.

Server configuration comes from command-line flags, a YAML file, or environment variables. The ones worth recognising are CHROMA_PERSIST_PATH (where data lives), CHROMA_PORT (default 8000), CHROMA_LISTEN_ADDRESS and CHROMA_ALLOW_RESET. Nested keys use a double underscore, as in CHROMA_OPEN_TELEMETRY__ENDPOINT. If you find an old tutorial using variables such as CHROMA_SERVER_AUTHN_PROVIDER, ignore it: built-in authentication no longer exists.

Resetting is deliberately hard. client.reset() deletes everything, so it is disabled by default and raises Reset is disabled by config. For experiments you can enable it with Settings(allow_reset=True) in your own process, or allow_reset: true on a server. Never enable it anywhere that holds data you care about.

Housekeeping. There is no built-in backup command for a single node. To back up, stop the server (or snapshot the volume) and copy the whole data folder, which holds chroma.sqlite3 and the index directories. There is also chroma vacuum --path <dir>, which compacts the SQLite file and is needed once if you upgrade a database created before version 0.5.6; it blocks reads and writes, so stop the server first. A note for readers of the official docs: their vacuum page shows chroma utils vacuum, but on 1.5.9 that fails with unrecognized subcommand 'utils'; the working form is chroma vacuum.

Try it
  1. Create a collection with configuration={"hnsw": {"space": "cosine"}}.
  2. Try to modify the configuration to {"hnsw": {"space": "l2"}} and read the error.
  3. Print the collection's configuration to confirm the space you chose.
a rejected change that proves the distance is locked at creation.

Common errors and how to read them

Chroma's errors are usually specific, which is a gift. Read the whole message; it generally names the exact value that was wrong. Here are the ones you will meet, grouped by what you were doing.

Creating and finding collections. Collection [demo] already exists means create_collection hit a name that is taken; use get_or_create_collection. Collection [nope] does not exist means the opposite: check the spelling, and check that you are connected to the right client, since an ephemeral client and a persistent client have different contents. Validation error: name: Expected a name containing 3-512 characters from [a-zA-Z0-9._-]... means the name broke the naming rules, often because it was too short.

Writing. Expected ID to be a str, got 1 means you used a number as an id; wrap it in str(). Unequal lengths for fields: ids: 2, embeddings: 1 means your parallel lists do not line up. Expected IDs to be unique, found duplicates of: z in add means one batch contained the same id twice. Expected metadata value to be a str, int, float, bool, SparseVector, list, or None, got {...} which is a dict means nested metadata. Expected metadata list value for key 't' to be non-empty means an empty list in metadata. A None metadata value on a local add produces Cannot convert Python object to MetadataValue; drop keys that are None. And Batch size 5462 exceeds maximum batch size 5461 means chunk your writes. Remember also the silent cases: adding an existing id and updating a missing id raise nothing.

Dimensions. Collection expecting embedding with dimension of 3, got 2 is almost always a mixed-model problem: your stored vectors came from one model and your query vectors from another. Use the same embedding function for both, and when you change models, make a new collection. Non-finite float value at embedding index 0: NaN means a vector contains a not-a-number; clean your data before sending it.

Querying. Number of requested results 0, cannot be negative, or zero means n_results=0. The include error about got ids means you listed ids in include. The where errors, empty filter or unknown operator, were covered earlier. You must provide an embedding function to compute embeddings means you called query_texts on a collection with no embedding function.

Connecting. Could not connect to a Chroma server means the server is down, the port is wrong, or Docker's -p is missing. Tenant [nope] not found and Database [nope] not found mean you passed names that do not exist; the defaults are default_tenant and default_database. An HTTP 410 on /api/v1 means an old client against a 1.x server; upgrade the client.

When stuck, check the three basics first Which client am I using, which collection name did I pass, and which embedding function does this collection have? Printing client.list_collections() and col.count() resolves a surprising share of "my data disappeared" problems in under a minute.
Try it
  1. Deliberately trigger four errors: a duplicate collection name, a numeric id, a wrong-length query vector and an empty where.
  2. For each, write down in one line what the message told you and how you would fix it.
the habit of reading the message before searching the internet. Most Chroma errors name the exact problem.

Putting it all together

Let us finish with one small end-to-end project that uses everything: a command-line tool that loads a text file of notes into a persistent collection and answers questions from it. It is deliberately short, but it is the real shape of a retrieval pipeline.

Create notes.txt with one note per line, then kb.py:

kb.py
import sys
import chromadb

client = chromadb.PersistentClient(path="./chroma")
col = client.get_or_create_collection(
    name="notes",
    configuration={"hnsw": {"space": "cosine"}},
)


def load(path: str) -> None:
    with open(path, encoding="utf-8") as f:
        lines = [line.strip() for line in f if line.strip()]
    col.upsert(
        ids=[f"note-{i}" for i in range(len(lines))],
        documents=lines,
        metadatas=[{"source": path, "line": i} for i in range(len(lines))],
    )
    print(f"loaded {len(lines)} notes, collection has {col.count()}")


def ask(question: str) -> None:
    res = col.query(query_texts=[question], n_results=3)
    for doc, dist in zip(res["documents"][0], res["distances"][0]):
        print(f"{dist:.3f}  {doc}")


if __name__ == "__main__":
    command, argument = sys.argv[1], sys.argv[2]
    load(argument) if command == "load" else ask(argument)

Use it like this:

BASH
python kb.py load notes.txt
python kb.py ask "how do I get my money back"

Walk through the design choices, because each one comes from earlier in this guide. It uses a persistent client so the data survives between commands. It uses get_or_create_collection so the script is safe to run repeatedly, and cosine distance, set once at creation. It loads with upsert and stable ids built from the line number, so re-running load never duplicates records. It stores metadata about where each note came from, so that later you can filter by source or show a citation. And it prints the distances next to each answer, so you can see how confident the match is.

To turn this into the retrieval half of a question-answering system, take the top documents from ask, put them in a prompt as context, and send that prompt to a language model; the LangChain guide and the LlamaIndex guide show frameworks that wrap exactly this pattern, and the Claude API guide covers the model call. If you would rather compare Chroma with other stores before choosing, the Qdrant guide and the pgvector guide cover two common alternatives.

One real-world limit to plan for early: a note long enough to contain several ideas gets one embedding that blurs them all together. Real systems chunk long documents into passages of a few sentences each before adding them. Keep each record about one idea, and your retrieval improves more than any setting change will.

Try it
  1. Write fifteen notes about a topic you know well and run the load command.
  2. Ask five questions in your own words, avoiding the exact words in the notes.
  3. Mark which ones returned the right note at rank one, and for the misses, decide whether the note or the question was the problem.
most questions answered correctly even with different wording, and a short list of misses you can reason about. That list is your first evaluation set.

What you can now do, and what comes next

You started with a database that finds text by meaning, and you can now use it for real. You can explain what an embedding is and why lower distance means a better match. You know the nouns, tenant, database, collection and record, and which ones you will touch at first. You can install and pin Chroma 1.5.9 and verify it. You can create ephemeral, persistent and server-backed clients and choose between them. You can add, update, upsert, delete, query and get, read the nested results, filter with where and where_document, supply your own vectors, run chroma run or Docker, and read the errors that come back.

Just as important, you know the quiet traps: adding an existing id does nothing, the distance is fixed at creation, a query always returns something, mixing embedding models breaks dimensions, and a self-hosted 1.x server has no login of its own.

What comes next in the mid-level guide: how to tune and measure retrieval quality, filter in depth, set HNSW parameters, write a custom embedding function, batch and parallelise loads, run with a configuration file, and test code that depends on Chroma. The senior guide covers architecture, sizing memory for the index, multi-tenancy, securing and upgrading a server, backups, observability, and when to choose something other than Chroma.

Practise before moving on. The best exercise is the one in the last section with your own data, because retrieval quality depends on your text far more than on any setting.

Sources