تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
WeaviateLLMsVector databases3 مستويات103 قسمًايغطّي Weaviate 1.39دليل بالإنجليزية

The Complete Weaviate Guide

Run hybrid vector and keyword search with the open-source Weaviate database. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
16sections
36examples

This is part one of three. It covers everything you need to start doing real work with Weaviate, not a teaser. By the end you can run a Weaviate server on your laptop, create a collection, load objects into it, search them by meaning, by keyword and by both at once, filter the results, generate answers from them, and read the common error messages without panic. Mid-level and Senior take the same topics further; nothing here is thrown away.

The guide targets Weaviate 1.39 (the stable line in August 2026, with patch releases continuing since) and the Python client v4. Weaviate also has clients for TypeScript, Go, Java and C#, and the ideas transfer directly, but every code sample here is Python.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas only stick once you have watched your own data go in and come back out.

What Weaviate is, and the problem it solves

Weaviate is an open-source vector database written in Go. It stores your data as objects, and next to each object it can store one or more vectors: long lists of numbers that capture what the object means. It then lets you search by meaning, by keyword, or by a blend of the two, and it can hand the results to a language model to produce an answer.

To see why that is useful, start with what came before it.

A traditional database is excellent at exact questions. "Give me the row where id = 42." "Give me every order placed after 1 March." If you want to find an article about getting a refund, a SQL LIKE '%refund%' query finds only the articles that contain that exact string. It misses the one titled "How to get your money back", even though a person would say it is the same topic. Keyword search engines improved on this with stemming and ranking, but they still match words, not ideas.

The change came from embedding models. An embedding model reads a piece of text (or an image, or audio) and returns a vector, for example a list of 768 floating-point numbers. The important property is that things with similar meaning get vectors that sit close together in that space. "Refund" and "money back" land near each other. "Refund" and "volcano" do not. Once your data is vectors, "find things similar in meaning to this query" becomes "find the vectors nearest to this query vector".

That nearest-neighbour question is expensive if you answer it naively. Comparing a query against ten million vectors one by one is far too slow. Vector databases exist to answer it fast, using an approximate nearest neighbour (ANN) index that trades a tiny amount of accuracy for a huge speed-up. They also do the unglamorous work around it: storing the original objects, filtering on ordinary fields, keeping data safe on disk, and scaling across machines.

Weaviate's own description is that it stores objects together with their vectors and serves vector, keyword (BM25), hybrid, filtered and generative search. The pieces you will meet in this guide are:

OBJECTSyour data
→
VECTORSmeaning as numbers
→
INDEXESfast lookup
→
SEARCHvector, keyword, hybrid
→
GENERATEoptional answers

The most common reason people reach for Weaviate is retrieval-augmented generation (RAG): you want a language model to answer questions about your own documents. You cannot paste ten thousand pages into a prompt. Instead you store the pages in a vector database, search for the few passages most relevant to the question, and give only those to the model. Weaviate does the storing and searching, and can even call the model for you. Other uses are semantic site search, recommendations ("more like this"), deduplication, and searching images by a text description.

A note on where Weaviate sits among its neighbours. It is not the only vector database: Qdrant and Pinecone are comparable dedicated systems, and pgvector adds vector search to PostgreSQL. Weaviate's character is that it is a full database with a schema, built-in hybrid search and optional built-in integrations with embedding and generation providers.

Try it
  1. Pick three sentences about different topics and one query sentence that shares a meaning (not a word) with one of them.
  2. Write down which sentence you expect a vector search to rank first, and which one a plain keyword search would rank first.
  3. Keep the sentences. You will load them into Weaviate in a later section and compare the result with your prediction.

The mental model: four nouns

Weaviate has more features than you will learn in a week, but the vocabulary for everything you do as a beginner is four words. Learn these and the rest of the documentation becomes readable.

Collection. A collection is a named set of objects that share a schema, much like a table in a relational database. You might have a collection called Article and another called Product. You will see the older word class in some documentation and in REST paths such as /v1/schema/{className}. It means the same thing as a collection. Collection names must start with an uppercase letter, so Article works and article does not.

Object. An object is one record in a collection. Each object has a UUID (a unique identifier, which Weaviate generates for you unless you supply one), a set of properties, zero or more vectors, and creation and update timestamps. If a collection is a table, an object is a row.

Property. A property is a typed field on an object, like a column. Common types are text, int, number (a floating-point number), boolean, date, uuid and arrays such as text[]. Dates use RFC 3339 format, for example 1985-04-12T23:20:50.52Z. Property names start with a lowercase letter, and Weaviate lowercases the first character for you.

Vector. A vector is the list of numbers that represents an object's meaning. Where it comes from is the most important choice you make, and there are two options. Either you configure a vectorizer (a model provider integration such as text2vec-openai or text2vec-ollama) and Weaviate calls the embedding model for you whenever you insert an object or run a text query. Or you bring your own vectors: you compute them yourself and pass them in. Weaviate calls that second option self-provided vectors.

Two further words appear constantly and are worth defining now.

A vector index is the data structure that makes nearest-neighbour search fast. The default is HNSW (Hierarchical Navigable Small World), a graph held in memory. You do not need to understand its internals to begin. You only need to know it is what Weaviate builds so that a search over a million objects takes milliseconds, and that it lives in RAM, which is why memory is Weaviate's main cost at scale.

A distance metric defines "close". The default is cosine, which compares the direction of two vectors and ignores their length. Search results report a distance, and lower means closer. Newcomers often expect a "score" where higher is better, so get used to reading distance the other way round.

Collections are not tables in one important way A relational table gets its meaning from the columns you query. A Weaviate collection also carries a decision about vectors: which vectorizer, which index, which distance metric. Several of those decisions cannot be changed after the collection is created. That is why the next sections have you think before you create, not after.

Weaviate talks to you through three interfaces, and knowing which does what prevents a lot of confusion later. REST on port 8080 handles creating collections, inserting and fetching objects, and health checks. gRPC on port 50051 is the fast path the v4 Python client uses for searching and batch imports, and the v4 client requires it. GraphQL at /v1/graphql is the older query interface. As a beginner using the Python client you will rarely touch GraphQL directly.

Try it
  1. Sketch a collection for something you care about, for example `Recipe` with properties `title`, `ingredients`, `minutes` and `vegetarian`.
  2. Choose a type for each property from the list above.
  3. Decide whether you would let Weaviate create the vectors from `title` and `ingredients`, or compute them yourself, and write one sentence on why.

Installing Weaviate and checking the setup

Weaviate ships as a Linux container image and as a Go binary. The easiest path on every operating system is Docker, so this guide uses it. If Docker is new to you, read the Docker guide first; the beginner part is enough. On Linux you need Docker Engine and the Compose plugin. On macOS, Docker Desktop, Colima or OrbStack all work, and Apple Silicon is supported through the arm64 image. On Windows use Docker Desktop with the WSL 2 backend.

There are other routes you should know exist. Weaviate also runs on Kubernetes through a Helm chart (weaviate/weaviate), which is the production route covered at the Mid and Senior levels. Weaviate Cloud is the managed service, where you create a cluster in a browser and receive a URL and an API key. And there is an embedded mode for Python that starts a Weaviate process from inside your script. Embedded mode is experimental and works only on Linux and macOS, so it is not a good teaching path for a mixed class.

Start the server

Create a folder for the project and put this in a file named docker-compose.yml. It is a single-node setup with data stored in a named volume:

docker-compose.yml
services:
  weaviate:
    command: [--host, 0.0.0.0, --port, '8080', --scheme, http]
    image: cr.weaviate.io/semitechnologies/weaviate:1.39.7
    ports:
    - 8080:8080
    - 50051:50051
    volumes:
    - weaviate_data:/var/lib/weaviate
    restart: on-failure:0
    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
      PERSISTENCE_DATA_PATH: '/var/lib/weaviate'
      CLUSTER_HOSTNAME: 'node1'
volumes:
  weaviate_data:

Read it line by line, because every line is something you will eventually need to change. The image pins an exact version, which is good practice: 1.39.7 was the latest 1.39 patch when this guide was verified, and you should prefer the newest patch of the minor you choose. The two ports lines publish the REST port 8080 and the gRPC port 50051 to your machine. Publishing 50051 is not optional, because the Python client checks it at startup. The volumes entry stores data in a Docker volume mounted at /var/lib/weaviate, which is where Weaviate writes. AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true' lets you connect without a key, which is convenient on a laptop and wrong for anything reachable from a network. QUERY_DEFAULTS_LIMIT sets how many results a query returns when you do not say.

Start it:

BASH
docker compose up -d

If you only want a quick throwaway run without a compose file, this also works:

BASH
docker run -p 8080:8080 -p 50051:50051 cr.weaviate.io/semitechnologies/weaviate:1.39.7

Without a volume, the data disappears when that container is removed. That is fine for a first experiment and a surprise if you forget it.

Check that it is alive

Weaviate has no core command-line tool. You check it with ordinary HTTP requests. The first tells you the version and the enabled modules:

BASH
curl http://localhost:8080/v1/meta

You get a JSON document that includes the version string and a list of modules. The readiness probe is the one to use in scripts:

BASH
curl -i http://localhost:8080/v1/.well-known/ready

A 200 OK means the node is ready to serve. There is also /v1/.well-known/live, which returns 200 as soon as the process is alive, even if it is still starting up. Right after docker compose up -d you may see a connection refused or a non-200 reply for a few seconds. Wait and retry; it is not an error.

Install the Python client

Weaviate's current Python client is weaviate-client, version 4.23 at the time of verification. It needs Python 3.10 or newer. Create a virtual environment and install it:

BASH
python -m venv .venv
source .venv/bin/activate
pip install -U weaviate-client

On Windows PowerShell, activate with .venv\Scripts\Activate.ps1 instead. Now confirm Python can reach the server:

check.py
import weaviate

client = weaviate.connect_to_local()
try:
    print(client.is_ready())
    print(client.get_meta()["version"])
finally:
    client.close()

You should see True and then a version such as 1.39.7. The try/finally is not decoration. The v4 client holds open connections, and you must close it. The tidier pattern is a with block, which closes the client for you:

PYTHON
import weaviate

with weaviate.connect_to_local() as client:
    print(client.is_ready())
Do not follow old tutorials Plenty of blog posts and videos still show the **version 3** client: weaviate.Client(...), client.schema.create_class(...), client.query.get(...).with_near_text(...). That client is no longer supported. If you see those names, the tutorial is out of date. The current entry points are weaviate.connect_to_*() and client.collections.*, which is what this guide uses.
Try it
  1. Run docker compose up -d, then docker compose ps and confirm the container is running.
  2. Run both curl checks above and find the version field in the /v1/meta output.
  3. Run check.py. Then stop the container with docker compose stop, run it again, and read the error you get. You will recognise it later.

Connecting from Python

Before building anything, spend a few minutes on the connection, because most beginner problems show up here. A connection function returns a client object, and everything else is a method on it.

The simplest is weaviate.connect_to_local(). It assumes Weaviate is at localhost with REST on 8080 and gRPC on 50051, which is what the compose file above provides.

When your setup differs, use connect_to_custom and spell out both channels:

PYTHON
import weaviate

client = weaviate.connect_to_custom(
    http_host="localhost", http_port=8080, http_secure=False,
    grpc_host="localhost", grpc_port=50051, grpc_secure=False,
)

There are two channels because the client talks over REST for some operations and gRPC for others. That is why a half-working setup is possible: REST reachable and gRPC blocked produces the most common connection error, covered in the errors section.

For Weaviate Cloud you pass the cluster URL and an API key. Read them from environment variables rather than writing them into a file you might commit:

PYTHON
import os
import weaviate
from weaviate.classes.init import Auth

client = weaviate.connect_to_weaviate_cloud(
    cluster_url=os.environ["WEAVIATE_URL"],
    auth_credentials=Auth.api_key(os.environ["WEAVIATE_API_KEY"]),
)

If you later turn on API-key authentication on your own server, the same Auth.api_key(...) object goes into connect_to_local through its auth_credentials argument.

Many embedding and generation providers need their own key, for example an OpenAI key. You hand it to Weaviate as a request header named X-<Provider>-Api-Key, and the client has a headers argument for it:

PYTHON
client = weaviate.connect_to_local(
    headers={"X-OpenAI-Api-Key": os.environ["OPENAI_API_KEY"]}
)

The alternative is to set the server's own environment variable (OPENAI_APIKEY), in which case every client shares one key. If both are present, the header wins. Provider keys are never stored in the collection configuration, which is a good thing: a backup of your database does not contain them.

Two last connection habits. Always close the client, either with with or client.close(); forgetting produces a ResourceWarning, and using a client after closing raises WeaviateClosedClientError. And if you are building a web service, create one client at startup and reuse it rather than connecting on every request.

Try it
  1. Rewrite check.py to use connect_to_custom with explicit hosts and ports.
  2. Change grpc_port to 50052, run it, and read the error message. Put it back.
  3. Check that client.is_live() and client.is_ready() both return True on a healthy server.

Your first collection and your first objects

Time to build something. We will start with the path that needs no API key and no extra services: bringing our own vectors. It is the best way to learn because nothing is hidden. Weaviate does not call any model. You see exactly what goes in and what comes out.

Real embedding vectors have hundreds of dimensions, which are unpleasant to type. For learning, we will use tiny three-dimensional vectors that we invent, where the three numbers loosely mean "fruit-ness", "vehicle-ness" and "animal-ness". The behaviour is identical to the real thing.

first_collection.py
import weaviate
from weaviate.classes.config import Configure, Property, DataType

with weaviate.connect_to_local() as client:
    if client.collections.exists("Thing"):
        client.collections.delete("Thing")

    things = client.collections.create(
        "Thing",
        vector_config=Configure.Vectors.self_provided(),
        properties=[
            Property(name="name", data_type=DataType.TEXT),
            Property(name="kind", data_type=DataType.TEXT),
            Property(name="weight_kg", data_type=DataType.NUMBER),
        ],
    )

    things.data.insert(
        properties={"name": "apple", "kind": "fruit", "weight_kg": 0.2},
        vector=[0.9, 0.1, 0.0],
    )
    things.data.insert(
        properties={"name": "banana", "kind": "fruit", "weight_kg": 0.12},
        vector=[0.8, 0.0, 0.1],
    )
    things.data.insert(
        properties={"name": "truck", "kind": "vehicle", "weight_kg": 9000},
        vector=[0.0, 0.95, 0.05],
    )
    things.data.insert(
        properties={"name": "horse", "kind": "animal", "weight_kg": 500},
        vector=[0.1, 0.1, 0.9],
    )

    print(things.aggregate.over_all(total_count=True).total_count)

Take it apart. client.collections.create("Thing", ...) creates a collection and returns a handle to it. The name starts with an uppercase letter, as required. vector_config=Configure.Vectors.self_provided() says "I will give you the vectors; do not call any model". The properties list declares three typed fields. data.insert(properties=..., vector=...) adds one object and returns its UUID. The last line asks the collection how many objects it holds; it should print 4.

Notice the guard at the top: if client.collections.exists("Thing"). Creating a collection that already exists fails with an error such as class name Thing already exists (HTTP 422). Making scripts re-runnable like this saves you from the most boring error in the guide.

Now search. Because we supplied vectors, we search with a vector too, using near_vector:

search_vectors.py
import weaviate
from weaviate.classes.query import MetadataQuery

with weaviate.connect_to_local() as client:
    things = client.collections.use("Thing")

    result = things.query.near_vector(
        near_vector=[0.85, 0.05, 0.05],   # "fruit-ish"
        limit=2,
        return_metadata=MetadataQuery(distance=True),
    )

    for obj in result.objects:
        print(obj.properties["name"], round(obj.metadata.distance, 4))

Expected output is the two fruits, apple and banana, with small distances, and the closer one first. client.collections.use("Thing") gets a handle to an existing collection. (You may meet client.collections.get("Thing") in older v4 examples. It was replaced by use, so prefer use.) The return_metadata argument asks for the distance; without it Weaviate returns only the properties. Each result object has .uuid, .properties and .metadata.

Look at the distances. They are small for the fruits and would be large for the truck. That is the entire idea of a vector database in one result: the query had no word in common with anything, and the right objects still came back, because the numbers were close.

Start with your own vectors when learning Using self_provided vectors removes three things that confuse beginners at once: API keys, network calls to a model, and cost. Once you are comfortable with collections, objects and queries, adding a vectorizer in the next section is a small step.
Try it
  1. Run first_collection.py, then search_vectors.py.
  2. Change the query vector to something vehicle-ish such as [0.0, 0.9, 0.1] and confirm the truck comes first.
  3. Add two more objects of your own with sensible vectors and check that searches still rank them sensibly.
  4. Run first_collection.py twice and notice that the guard makes the second run safe.

Letting Weaviate create the vectors

Typing vectors by hand teaches the idea, but in real projects you do not do it. You let an embedding model produce them. Weaviate supports this through modules (also called model provider integrations): you name a vectorizer when you create the collection, and from then on Weaviate calls the model whenever it needs a vector, both when you insert an object and when you run a text query.

The advantage is that your application code stays text-only. You insert {"title": "...", "body": "..."} and later search with near_text(query="..."). You never see a vector.

There are many providers: OpenAI, Cohere, Google, AWS, Mistral, Voyage AI, Hugging Face, Weaviate's own hosted embeddings (text2vec-weaviate) and more. For a course, a local model is attractive because it needs no account and no per-call cost. Weaviate's official local quickstart uses Ollama, a tool that runs open models on your own machine (see the Ollama guide for more).

The setup adds an Ollama container to the compose file and turns on two Weaviate modules, one for embeddings and one for generation. Replace docker-compose.yml with:

docker-compose.yml
services:
  weaviate:
    command: [--host, 0.0.0.0, --port, '8080', --scheme, http]
    image: cr.weaviate.io/semitechnologies/weaviate:1.39.7
    ports:
    - 8080:8080
    - 50051:50051
    volumes:
    - weaviate_data:/var/lib/weaviate
    restart: on-failure:0
    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'
      PERSISTENCE_DATA_PATH: '/var/lib/weaviate'
      ENABLE_MODULES: 'text2vec-ollama,generative-ollama'
      CLUSTER_HOSTNAME: 'node1'
  ollama:
    image: ollama/ollama:0.12.9
    ports:
    - 11434:11434
    volumes:
    - ollama_data:/root/.ollama
volumes:
  weaviate_data:
  ollama_data:

Apply it and download two models: an embedding model, nomic-embed-text, and a small chat model, llama3.2, which we use later for generation:

BASH
docker compose up -d
docker compose exec ollama ollama pull nomic-embed-text
docker compose exec ollama ollama pull llama3.2

The downloads are a few gigabytes in total, so do this on a good connection. Now the part that trips almost everyone, so read it twice.

Inside Docker, localhost is not your laptop Weaviate runs in its own container. From inside it, localhost means that container, where nothing listens on port 11434. Weaviate must reach Ollama by its compose service name: http://ollama:11434. Your Python script, which runs on your laptop, still uses localhost for Weaviate. Mixing these up gives "connection refused" from Weaviate.

Create a collection that uses Ollama to vectorize text:

ollama_collection.py
import weaviate
from weaviate.classes.config import Configure, Property, DataType

with weaviate.connect_to_local() as client:
    if client.collections.exists("Article"):
        client.collections.delete("Article")

    client.collections.create(
        "Article",
        vector_config=Configure.Vectors.text2vec_ollama(
            api_endpoint="http://ollama:11434",
            model="nomic-embed-text",
        ),
        generative_config=Configure.Generative.ollama(
            api_endpoint="http://ollama:11434",
            model="llama3.2",
        ),
        properties=[
            Property(name="title", data_type=DataType.TEXT),
            Property(name="body", data_type=DataType.TEXT),
        ],
    )

Two new arguments appear. vector_config=Configure.Vectors.text2vec_ollama(...) is the vectorizer: it names the endpoint and the embedding model. generative_config names the model Weaviate will call when you ask it to generate text. Both point at http://ollama:11434, as the warning explained.

A note on naming that will save you time. Older tutorials write vectorizer_config=Configure.Vectorizer.text2vec_openai() or Configure.NamedVectors.... Those forms are deprecated. The current form is vector_config=Configure.Vectors.<provider>(...), as above. If you prefer a hosted provider, swap in Configure.Vectors.text2vec_openai() and pass your key as a header, as shown in the connection section.

By default, Weaviate builds the vector from all text properties of the object, joined together. You can narrow it with source_properties, which the Mid-level guide covers. For now the default is sensible.

Try it
  1. Bring up the new compose file and pull both models.
  2. Run docker compose exec ollama ollama list and confirm both models are present.
  3. Run ollama_collection.py. Then call client.collections.list_all() and confirm Article exists.
  4. Deliberately change api_endpoint to http://localhost:11434, try inserting an object in the next section, and read the failure. Then fix it.

Loading data: single inserts and imports

With a vectorizer in place, loading data is just handing over properties. Weaviate calls the embedding model, stores the vector and indexes it.

load_articles.py
import weaviate

articles_data = [
    {"title": "How to get your money back",
     "body": "Steps for requesting a refund after a purchase, and how long it takes."},
    {"title": "Training a neural network on a laptop",
     "body": "Tips for fitting small models into limited memory and keeping GPUs cool."},
    {"title": "A beginner's guide to sourdough",
     "body": "Flour, water, salt and patience: how to keep a starter alive."},
    {"title": "Reading a bank statement",
     "body": "What each line means, and how to spot a charge you did not make."},
]

with weaviate.connect_to_local() as client:
    articles = client.collections.use("Article")

    first_id = articles.data.insert(articles_data[0])
    print("inserted", first_id)

    articles.data.ingest(articles_data[1:])
    print(articles.aggregate.over_all(total_count=True).total_count)

Two methods, two purposes. data.insert(properties) adds one object and returns its UUID. It is fine for a handful of objects or for an application saving one record at a time. data.ingest(list_of_objects) is the bulk path. It streams a whole list to the server using server-side batching, so you do not tune batch sizes yourself. It is the one to use for anything beyond a few dozen objects. The line at the end should print 4.

Why does bulk loading deserve its own method? Every object needs an embedding, and embedding is the slow, expensive step. Sending objects one network call at a time is slow, and sending them all in one giant request can be rejected for being too large. The gRPC channel has a message size cap (GRPC_MAX_MESSAGE_SIZE, about 100 MB by default), and a huge payload fails with Sent message larger than max. ingest avoids both problems. You may see an older method insert_many in tutorials. It works for small lists, but it is the one that triggers that size error on big ones.

Three details about objects are worth knowing from day one.

First, UUIDs. If you do not supply an ID, Weaviate generates a random one. Running the loading script twice then gives you duplicates, because every run creates new random IDs. If you want re-runs to be safe, derive the ID from the content with the helper generate_uuid5:

PYTHON
from weaviate.util import generate_uuid5

obj_id = generate_uuid5(articles_data[0])
articles.data.insert(articles_data[0], uuid=obj_id)

The same content always yields the same UUID, so inserting it again fails with an "already exists" error instead of silently adding a duplicate. That is a feature: it turns a blind duplicate into a visible message.

Second, autoschema. If you insert a property that the collection does not define, Weaviate by default creates it for you, guessing the type from the value. That is AUTOSCHEMA_ENABLED, on by default. It is handy while exploring, but it also means a typo in a property name quietly creates a new property. Declare your properties explicitly, as we did, and turn autoschema off on servers that matter.

Third, errors in bulk loads do not always raise. A batch can partly succeed. After a batched load, check for failures rather than assuming success. Use the batch.stream() context manager when you want finer control and look at articles.batch.failed_objects afterwards. The Mid-level guide covers this in depth. For the beginner, the habit is: after importing, count the objects and compare with what you sent.

Count after you load articles.aggregate.over_all(total_count=True).total_count is the quickest sanity check in Weaviate. If you sent 1,000 objects and the count says 940, something failed, and you want to know now rather than when a search misses an obvious result.
Try it
  1. Run load_articles.py and confirm the count is 4.
  2. Run the load again and notice the count rise to 8, because random UUIDs gave you duplicates.
  3. Delete the collection by re-running ollama_collection.py, then reload using generate_uuid5 IDs and see how a second run behaves.

Searching: by meaning, by keyword, and by both

This is the payoff. Weaviate offers several ways to search, and choosing between them is the main skill of a beginner. All of them are methods on collection.query.

Search by meaning with near_text

near_text sends your query text to the vectorizer, gets a vector, and finds the nearest objects:

search_text.py
import weaviate
from weaviate.classes.query import MetadataQuery

with weaviate.connect_to_local() as client:
    articles = client.collections.use("Article")

    result = articles.query.near_text(
        query="I want to be reimbursed",
        limit=2,
        return_metadata=MetadataQuery(distance=True),
    )

    for obj in result.objects:
        print(round(obj.metadata.distance, 3), obj.properties["title"])

The query contains none of the words "refund", "money" or "back", and the refund article should still come first. That is semantic search. Exact distances depend on the embedding model, so your numbers will differ from anyone else's. What matters is the order.

Search by keyword with bm25

bm25 is classic keyword search. It scores objects by how often the query's words appear, weighting rare words more heavily. No embedding model is involved:

PYTHON
result = articles.query.bm25(query="sourdough starter", limit=3)
for obj in result.objects:
    print(obj.properties["title"])

Keyword search wins when the user types something specific that must appear literally: a product code, an error message, a person's name. Vector search can blur such exact tokens. If the query is "ERR_4021", you want BM25.

Combine them with hybrid

hybrid runs both a vector search and a BM25 search and fuses the two rankings:

PYTHON
result = articles.query.hybrid(query="refund policy", alpha=0.5, limit=3)
for obj in result.objects:
    print(obj.properties["title"])

The alpha parameter sets the balance: 1 means pure vector search, 0 means pure keyword search, and 0.5 weighs them equally. By default the two result lists are fused with relative score fusion, the method Weaviate has used for hybrid search since version 1.24. The other option is rank-based fusion (HybridFusion.RANKED), a topic for later.

Which should a beginner choose? A reasonable default for text search in real applications is hybrid, because it handles both the vague question and the exact term. Start with alpha=0.5 and adjust after looking at real results. Use pure near_text when queries are always natural language, and pure bm25 when they are always exact tokens.

Search is not always about text

You have already used near_vector with your own vectors. Weaviate also has near_object (find objects similar to a given object, which is the "more like this" feature) and, with a multimodal module, near_image. They all follow the same shape: a query, a limit, and optional metadata.

Reading results

Every search returns an object with an .objects list. Each item has:

  • .uuid, the object's identifier;
  • .properties, a dictionary of the stored fields;
  • .metadata, which holds distance, score or other values only if you asked for them with return_metadata=MetadataQuery(...).

For bm25 and hybrid, the relevant metadata is score (higher is better), requested with MetadataQuery(score=True). For vector searches it is distance (lower is better). Mixing these up is the classic beginner mistake of reading a distance as a score.

A last note about limit. If you omit it, the server applies its default (QUERY_DEFAULTS_LIMIT, which our compose file set to 25). Always set limit explicitly so your code does not depend on server configuration.

Vector search (near_text)

  • Finds meaning, not words
  • Handles paraphrase and synonyms
  • Needs an embedding model
  • Can blur exact codes and names

Keyword search (bm25)

  • Matches the literal words
  • Excellent for IDs, codes, names
  • No model needed
  • Misses "money back" for "refund"
Try it
  1. Run the same query, "I want to be reimbursed", with near_text, bm25 and hybrid. Write down the top result of each.
  2. Compare with the prediction you made in the first Try it. Was the vector search right?
  3. Run hybrid with alpha=0 and alpha=1 and confirm they behave like bm25 and near_text respectively.

Filtering: combining meaning with exact conditions

Real questions mix fuzzy and exact: "articles about refunds, written this year, in English". Vector similarity cannot express "this year". For exact conditions you use filters. A filter limits which objects are even considered, and you attach it to any search.

Filters use the Filter class. Our Article collection has only text properties, so to demonstrate, add a richer collection. This one uses self-provided vectors again, so it works with no models running:

products.py
import weaviate
from weaviate.classes.config import Configure, Property, DataType
from weaviate.classes.query import Filter

with weaviate.connect_to_local() as client:
    if client.collections.exists("Product"):
        client.collections.delete("Product")

    products = client.collections.create(
        "Product",
        vector_config=Configure.Vectors.self_provided(),
        properties=[
            Property(name="name", data_type=DataType.TEXT),
            Property(name="category", data_type=DataType.TEXT),
            Property(name="price", data_type=DataType.NUMBER),
            Property(name="in_stock", data_type=DataType.BOOL),
        ],
    )

    products.data.insert({"name": "Trail shoes", "category": "outdoor",
                          "price": 89.0, "in_stock": True}, vector=[0.9, 0.1])
    products.data.insert({"name": "Hiking boots", "category": "outdoor",
                          "price": 140.0, "in_stock": False}, vector=[0.85, 0.15])
    products.data.insert({"name": "Office chair", "category": "furniture",
                          "price": 210.0, "in_stock": True}, vector=[0.1, 0.9])

    cheap_outdoor = products.query.near_vector(
        near_vector=[0.9, 0.1],
        limit=5,
        filters=(
            Filter.by_property("category").equal("outdoor")
            & Filter.by_property("price").less_than(100)
        ),
    )
    for obj in cheap_outdoor.objects:
        print(obj.properties["name"], obj.properties["price"])

The boolean type is spelled DataType.BOOL in the Python client. Only the trail shoes pass both conditions. The pieces: Filter.by_property("category") selects a property, and an operator method compares it. The available operators include equal, not_equal, less_than, less_or_equal, greater_than, greater_or_equal, like (with * wildcards, for example like("*shoe*")), and for array properties contains_any and contains_all. Combine conditions with & (and) and | (or), putting each condition in parentheses when you chain several.

You can also use a filter without any vector search, with fetch_objects. That is Weaviate acting like an ordinary database:

PYTHON
result = products.query.fetch_objects(
    filters=Filter.by_property("in_stock").equal(True),
    limit=10,
)
for obj in result.objects:
    print(obj.properties["name"])

There is also fetch_object_by_id(uuid) to get a single object when you know its ID.

One subtlety trips beginners with text. Filters and keyword matching work on tokens, not on the whole string. By default text is tokenized by word, which splits it into lowercase words. So an equal("outdoor") match on a multi-word text property matches objects that contain that word, and a filter on a full sentence behaves differently from what you might expect. When you need an exact whole-value match, such as an ID stored as text, create the property with tokenization=Tokenization.FIELD (imported from weaviate.classes.config). The Mid-level guide returns to this.

Try it
  1. Run products.py and confirm only the trail shoes are returned.
  2. Change the filter to in_stock == True and category == "outdoor" and predict the result before running.
  3. Use fetch_objects with Filter.by_property("price").greater_than(100) and list the names.

Generative search: asking a model about your results

So far Weaviate found objects. Generative search goes one step further: it finds objects and then asks a language model to do something with them. This is RAG in a single call. It uses the generate namespace, and it needs the generative_config we gave the Article collection.

rag.py
import weaviate

with weaviate.connect_to_local() as client:
    articles = client.collections.use("Article")

    response = articles.generate.near_text(
        query="problems with my bank account",
        limit=2,
        single_prompt="Summarise this article in one sentence: {body}",
        grouped_task="In two sentences, say what these articles have in common.",
    )

    for obj in response.objects:
        print(obj.properties["title"])
        print("  ->", obj.generative.text)

    print("Overall:", response.generative.text)

There are two ways to prompt and you can use either or both.

single_prompt runs the model once per retrieved object. Curly braces interpolate that object's properties, so {body} is replaced with each article's body in turn. The result for each object is at obj.generative.text.

grouped_task runs the model once over all the results together, and the answer is at response.generative.text. Use it for summaries and questions across the whole result set.

The first call is slow, because the local model has to load into memory. If the model name is wrong, or Ollama is not reachable at http://ollama:11434, you get an error from Weaviate about the generative module, not from Python.

You can override the generation model for a single query without changing the collection. Version 1.30 added that:

PYTHON
from weaviate.classes.generate import GenerativeConfig

response = articles.generate.near_text(
    query="problems with my bank account",
    limit=2,
    grouped_task="Name the main theme.",
    generative_provider=GenerativeConfig.ollama(
        api_endpoint="http://ollama:11434",
        model="llama3.2",
    ),
)
print(response.generative.text)

The same generate namespace exists for bm25, hybrid and near_vector, so you can retrieve with whichever search suits the data and then generate.

Keep expectations honest. Generation quality depends on the model you choose and on how good the retrieved objects are. A small local model will be weaker than a large hosted one. And if retrieval returns the wrong passages, the answer will be wrong, however good the model. Most "the RAG gives bad answers" problems are search problems, which is why the quality of your search is worth more attention than the prompt. The RAGAS guide covers how to measure that.

Try it
  1. Run rag.py and read both the per-object and the grouped output.
  2. Change grouped_task to ask a question your articles can answer, such as "How long does a refund take?", and check whether the answer comes from the article text.
  3. Ask a question the articles cannot answer and note what the model does. This shows why retrieval quality matters.

Everyday operations: update, delete, inspect

Once data is in, you need to change it, remove it and look around. These are the operations you will use constantly.

Update an object. data.update merges the properties you give into the existing object, leaving other properties alone. data.replace overwrites the object entirely:

PYTHON
articles.data.update(uuid=first_id, properties={"title": "How to get a refund"})

When the text of a vectorized property changes, Weaviate recomputes the vector. A merge that changes a vectorized property therefore costs another embedding call.

Delete an object. By ID:

PYTHON
articles.data.delete_by_id(first_id)

Delete many objects by filter uses delete_many:

PYTHON
from weaviate.classes.query import Filter

articles.data.delete_many(where=Filter.by_property("title").like("*sourdough*"))

Be careful: a delete filter that matches more than you intended deletes it all. Run the same filter through fetch_objects first and look at what it returns. Weaviate does reject a batch delete that omits its match condition, returning 422 Unprocessable Entity since 1.39 (it used to return a 500).

Read everything. For a full pass over a collection, use the iterator, which pages through objects with a cursor and has no result cap:

PYTHON
for obj in articles.iterator():
    print(obj.uuid, obj.properties["title"])

Ordinary queries are limited by QUERY_MAXIMUM_RESULTS (10,000 by default), so a search with a huge limit is not the way to export data.

Inspect collections.

PYTHON
print(client.collections.list_all())
print(client.collections.exists("Article"))

Delete a collection. This deletes the collection and all its objects:

PYTHON
client.collections.delete("Article")

There is no undo and no confirmation prompt. Treat it like DROP TABLE.

Add a property later. You can add a new property to an existing collection:

PYTHON
from weaviate.classes.config import Property, DataType

articles.config.add_property(Property(name="author", data_type=DataType.TEXT))

What you generally cannot do afterwards is change the settings that define the collection: a property's data type, the distance metric, the vectorizer. For those, you create a new collection and copy the data across. This is the strongest argument for the "think before you create" advice. A beginner who wants to change the embedding model is not stuck, but they are going to rebuild the collection, and re-embedding a big dataset takes time and money.

The REST equivalents. Everything above also exists over plain HTTP, which is handy for debugging with curl. GET /v1/schema lists collections and their configuration, GET /v1/objects lists objects, POST /v1/batch/objects batch-loads them and DELETE /v1/objects/{className}/{id} deletes one. You will use GET /v1/schema often to confirm what a collection really looks like:

BASH
curl http://localhost:8080/v1/schema
Try it
  1. Update one article's title, then fetch it by ID with fetch_object_by_id and confirm the change.
  2. Use iterator() to print every title in the collection.
  3. Run curl http://localhost:8080/v1/schema and find the vectorizer and your property names in the JSON.

Configuration you should know about

Weaviate is configured almost entirely through environment variables set on the server, which in our setup means the environment: block of the compose file. You do not need to memorise them. You need to know the handful that matter early and where to find the rest: the environment-variable reference in the official docs.

The ones that show up in a beginner's life:

Variable What it does
AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED Lets clients connect without a key. Anonymous access is on in the quickstart setups. Turn it off for anything shared
PERSISTENCE_DATA_PATH Where data lives inside the container. Mount a volume here
ENABLE_MODULES Which model integrations are available, such as text2vec-ollama,generative-ollama
QUERY_DEFAULTS_LIMIT How many results you get when you do not pass limit
QUERY_MAXIMUM_RESULTS The hard cap on results per query (10,000)
AUTOSCHEMA_ENABLED Auto-create properties on insert. On by default
GRPC_MAX_MESSAGE_SIZE The largest gRPC message accepted, about 100 MB
DISABLE_TELEMETRY Opts out of anonymous usage telemetry
LOG_LEVEL How chatty the server logs are

A word on modules. A module is an optional plug-in that adds a capability: a vectorizer, a generative model, a reranker, a backup target. Since version 1.33, modules that call hosted API providers are enabled by default, and the variable API_BASED_MODULES_DISABLED=true turns them off. Modules that run locally, like text2vec-ollama, you list in ENABLE_MODULES. The /v1/meta endpoint shows which are active.

Authentication deserves a clear statement, because the quickstart configuration hides it. The setup in this guide has anonymous access on: anyone who can reach port 8080 can read, change and delete everything. That is acceptable on localhost. The moment a server is reachable by anyone else, enable API-key authentication and authorization instead. The relevant variables are AUTHENTICATION_APIKEY_ENABLED, AUTHENTICATION_APIKEY_ALLOWED_KEYS and AUTHENTICATION_APIKEY_USERS, with role-based access control through AUTHORIZATION_RBAC_ENABLED and AUTHORIZATION_RBAC_ROOT_USERS. If you turn RBAC on without naming at least one root user, Weaviate refuses to start. This is a topic for the Mid-level guide. The beginner's rule is simply: never put an anonymous Weaviate on a public address.

Persistence is the other thing to get right. The volume at /var/lib/weaviate is where your data lives. docker compose down stops and removes containers but keeps the volume, so your data survives. docker compose down -v also deletes the volumes, and with them all your data. The -v is easy to type from muscle memory.

One flag deletes your database docker compose down -v removes the named volumes. If your Weaviate data was in weaviate_data, it is gone, with no recovery. Use plain docker compose down unless you truly want a fresh start.
Try it
  1. Run docker compose down, then docker compose up -d, and confirm your Article objects are still there.
  2. Open the /v1/meta output and find the list of enabled modules.
  3. Set QUERY_DEFAULTS_LIMIT to 3, restart, run a query with no limit, and count the results.

Reading common errors

Most Weaviate errors fall into a small set. Learn to recognise them by their wording.

WeaviateGRPCUnavailableError: gRPC health check could not be completed. The client reached the REST port but not the gRPC port. Almost always, port 50051 is not published, is blocked, or the client is pointed at the wrong gRPC host or grpc_secure setting. In Docker, make sure the compose file has 50051:50051. In Kubernetes, the gRPC service must be enabled in the Helm values. Use connect_to_custom to set the gRPC host and port explicitly. There is a skip_init_checks=True option, but treat it as a temporary diagnostic, not a fix.

class name Article already exists (HTTP 422). You tried to create a collection that is already there. Check client.collections.exists(...) first, or delete it.

401 Unauthorized, or a message about anonymous access not being enabled. The server requires authentication and your client sent nothing. Pass Auth.api_key(...) or enable anonymous access on a local dev server.

403 Forbidden. You authenticated, but your user's role lacks permission for that action. This is RBAC at work.

Vectorizer returns 401 or 429 on insert. The model provider rejected the call: a missing or wrong key (401), or you are sending requests faster than the provider allows (429). Check the X-<Provider>-Api-Key header or the server's <PROVIDER>_APIKEY variable, and slow your imports.

Ollama "connection refused". Weaviate is trying localhost from inside its container. Use http://ollama:11434.

Sent message larger than max. One request exceeded the gRPC size cap. Load with data.ingest(...) instead of one huge insert_many.

WeaviateClosedClientError, or a ResourceWarning about unclosed connections. You used a client after closing it, or never closed it. Use a with block.

The server will not accept writes and shards show READONLY. The disk or memory crossed a safety threshold (by default 90 percent disk use makes shards read-only). Free space and restore the shard's status. You will meet this in real use, rarely on a laptop.

All errors raised by the Python client inherit from weaviate.exceptions.WeaviateBaseError, with more specific classes such as WeaviateConnectionError, WeaviateQueryError and WeaviateBatchError. A try/except on the specific class lets you handle a connection failure differently from a bad query.

When an error is unclear, two sources of truth help. The container logs, from docker compose logs weaviate, show the server's side of the story, including module and vectorizer failures that the client reports only vaguely. And curl http://localhost:8080/v1/meta confirms the version and modules you are actually running, which settles many "but I configured it" disputes.

Try it
  1. Run docker compose logs weaviate and find the line that shows Weaviate started.
  2. Reproduce the "already exists" error by creating the same collection twice without the guard.
  3. Stop only the Ollama container with docker compose stop ollama, then try inserting an article. Read both the Python error and the Weaviate log line, then start Ollama again.

Putting it all together

Let us finish with one small end-to-end project that uses everything: a searchable help centre with a question-answering endpoint. It has three steps and one file each.

First, the setup script creates the collection with explicit properties, including one for filtering:

setup_helpcenter.py
import weaviate
from weaviate.classes.config import Configure, Property, DataType

with weaviate.connect_to_local() as client:
    if client.collections.exists("HelpArticle"):
        client.collections.delete("HelpArticle")

    client.collections.create(
        "HelpArticle",
        vector_config=Configure.Vectors.text2vec_ollama(
            api_endpoint="http://ollama:11434", model="nomic-embed-text"),
        generative_config=Configure.Generative.ollama(
            api_endpoint="http://ollama:11434", model="llama3.2"),
        properties=[
            Property(name="title", data_type=DataType.TEXT),
            Property(name="body", data_type=DataType.TEXT),
            Property(name="topic", data_type=DataType.TEXT),
        ],
    )

Second, the loader reads articles from a JSON file named articles.json (a list of objects with title, body and topic), derives stable IDs so re-runs do not duplicate, and verifies the count:

load_helpcenter.py
import json
import weaviate
from weaviate.classes.data import DataObject
from weaviate.util import generate_uuid5

with open("articles.json", encoding="utf-8") as f:
    rows = json.load(f)

with weaviate.connect_to_local() as client:
    col = client.collections.use("HelpArticle")
    col.data.insert_many([
        DataObject(properties=r, uuid=generate_uuid5(r["title"])) for r in rows
    ])
    print("expected", len(rows), "stored",
          col.aggregate.over_all(total_count=True).total_count)

For a few dozen articles insert_many is fine, and DataObject lets us attach a stable UUID to each. For thousands, switch to the batching approach from the Mid-level guide. Notice the last two lines: the loader compares what it sent with what is stored.

Third, the question-answering script retrieves with hybrid search, optionally limited to a topic, and has the model answer from the retrieved articles:

ask.py
import sys
import weaviate
from weaviate.classes.query import Filter, MetadataQuery

question = " ".join(sys.argv[1:]) or "How do I get a refund?"

with weaviate.connect_to_local() as client:
    col = client.collections.use("HelpArticle")

    result = col.generate.hybrid(
        query=question,
        alpha=0.5,
        limit=3,
        filters=Filter.by_property("topic").equal("billing"),
        return_metadata=MetadataQuery(score=True),
        grouped_task=f"Answer this question using only the articles provided: {question}",
    )

    print("Question:", question)
    for obj in result.objects:
        print(f"  {obj.metadata.score:.3f}  {obj.properties['title']}")
    print("Answer:", result.generative.text)

Run it with python ask.py "How long does a refund take?". Each step maps to a concept from this guide. The collection and properties are the schema. The vectorizer turns text into vectors automatically. hybrid combines meaning and keywords. The Filter restricts results with an exact condition. generate with grouped_task turns the retrieved objects into an answer, and the sources list above it shows which articles that answer rests on.

If you want to take it further, wrap ask.py in a small FastAPI service, or use a framework such as LangChain to orchestrate it. The database layer you have just learned stays the same.

Try it
  1. Write ten short help articles in articles.json, split between two topics, one of them billing.
  2. Run the three scripts in order and ask three questions: one clearly answered, one worded with synonyms, and one not covered at all.
  3. Remove the filters argument and see how the results change.

What you can now do, and what comes next

You can now run Weaviate in Docker, check that it is healthy, and connect to it from Python with the v4 client. You can create collections with typed properties, choose between bringing your own vectors and letting a module like Ollama create them, and load data with insert, ingest and stable UUIDs. You can search by meaning, keyword or both, narrow results with filters, and have a model generate answers over what was retrieved. You can update, delete and inspect data, read the main configuration variables, and diagnose the errors beginners meet most.

What you have not yet touched is what separates a demo from a service. Mid-level goes into batching with error handling, named vectors (several vectors per object), tuning hybrid search, filters on references and metadata, multi-tenancy for serving many customers from one collection, quantization to shrink memory, schema changes and aliases, and the async client. Senior covers replication and consistency, sharding, backups and restores, authentication with OIDC and RBAC, upgrades, monitoring with Prometheus, cost and capacity planning, and when Weaviate is the wrong tool.

Two practical pointers before you go. Weaviate releases often, and a new minor version appears every couple of months; only the latest three minor versions receive fixes, so check the release notes before you start a project and pin the exact version in your compose file. And look at how Weaviate compares with its neighbours in this catalogue, Qdrant, Milvus, Chroma and pgvector, so you can explain why you chose it.

For readers working with Gulf or Egyptian employers, one more practical point: where the vectors are computed matters for data residency. With a local model such as Ollama, document text never leaves your infrastructure. With a hosted embedding provider, the text of every object is sent to that provider's API, which may be an issue for regulated data. Weaviate lets you choose either, per collection, and that is worth raising early in a design discussion.

Sources