تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
FAISSLLMsVector databases3 مستويات111 قسمًايغطّي FAISS 1.15دليل بالإنجليزية

The Complete FAISS Guide

Build fast similarity-search indexes in-process with Meta’s FAISS library. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
18sections
24examples

This is part one of three. It covers everything you need to do real work with FAISS as a first-time user. By the end you can install it, build an index from your own vectors, search it, save it, load it again, make it faster with an IVF or HNSW index, read its most common error messages, and wire it to a small text-search project. Mid-level and Senior take the same topics further; nothing here is thrown away.

Each section ends with a Try it task. Do them as you go. They take a few minutes each, and the ideas in this guide only stick once you have watched your own search return a wrong answer, and then fixed it.

Everything below was checked against FAISS 1.15.1, the current release at the time of writing.

What FAISS is, and the problem it solves

FAISS stands for Facebook AI Similarity Search. It is an open-source library from Meta, written in C++ with Python bindings, that answers one question very quickly: given a vector, which stored vectors are closest to it?

To see why that question matters, start with what a vector is in this context. A vector (also called an embedding) is a list of numbers, for example 384 or 768 of them, produced by a machine-learning model. The useful property of embeddings is that things with similar meaning end up with similar numbers. A sentence-embedding model turns "How do I reset my password?" and "I forgot my login credentials" into two lists of numbers that sit close together, even though the two sentences share almost no words. The same trick works for images, audio clips, products, and code.

Once your data is vectors, "find things similar to this one" becomes "find the stored vectors closest to this query vector". That is the job of FAISS. It powers semantic search, recommendation, duplicate detection, and the retrieval half of retrieval-augmented generation (RAG), where a language model is given relevant documents found by vector search before it answers.

YOUR DATAtext, images, items
→
EMBEDDING MODELdata to vectors
→
FAISS INDEXstores and searches
→
NEAREST IDSrow numbers back

The diagram is the whole workflow. FAISS owns only the third box. It does not create embeddings and it does not store your documents. It stores vectors and hands back the row numbers of the closest ones, and you use those numbers to look up the original text, image, or product in your own storage.

Why does this need a dedicated library? Because the obvious approach, comparing the query against every stored vector, gets expensive fast. With a thousand vectors it is instant. With ten million vectors of 768 numbers each, one query means about seven and a half billion multiplications. FAISS is built to make that fast: it uses optimised linear algebra and CPU vector instructions for the exact approach, and it offers approximate index types that skip most of the data while still finding nearly all the right answers. It can also run on NVIDIA GPUs.

Two things set FAISS apart from other tools you may have heard of, and both explain most of the surprises that beginners hit.

FAISS is a library, not a database. There is no server to start, no port, no login, no query language. You import faiss in your own Python process, and the index lives in that process's memory. If your program exits, the index is gone unless you saved it to a file. Vector databases such as Qdrant, Milvus, Weaviate, Chroma and pgvector add storage, networking, metadata, and updates on top of ideas similar to those in FAISS, and a few of them use FAISS inside. With FAISS you get the engine and you build the car yourself.

FAISS only knows numbers. It stores float32 vectors and returns 64-bit integer ids. It has no idea what text, filenames, or categories are. Keeping the mapping between ids and your real data is your job. This is easy once you expect it and bewildering if you do not.

Try it
  1. Write down three searches you use every day (a music app, an online shop, a help centre). For each, decide whether it matches exact words or "similar meaning".
  2. For the "similar meaning" ones, write what the vector for a stored item would have to capture for the search to work.
  3. Estimate the work for brute force: a collection of 2 million items with 768-number vectors needs about how many multiplications per query? (Answer: 2,000,000 x 768, roughly 1.5 billion.)

Where FAISS fits, and when to pick something else

Before FAISS, nearest-neighbour search in Python usually meant one of three things: a NumPy loop or matrix multiplication, scikit-learn's nearest-neighbour classes, or a hand-rolled tree structure. These all work for small data. They slow down badly when the number of vectors reaches millions or the number of dimensions reaches hundreds, because tree structures stop helping in high dimensions and brute force simply costs too much.

FAISS gives you a menu of trade-offs. At one end is exact search, which is always right and gets slower as data grows. At the other end are approximate indexes, which answer in a small fraction of the time and memory but occasionally miss a true neighbour. The skill you will build in this guide is choosing a point on that menu on purpose rather than by accident.

Because FAISS is a library, you should also know when it is the wrong tool. Reach for a vector database instead when you need any of the following:

  • Updates and deletes at high frequency, from many users at once.
  • Filtering by metadata, such as "only documents from 2025 in the legal category". FAISS has no metadata filtering. It can filter only by integer id.
  • A service that several applications call over the network, with authentication.
  • Durability and replication that you do not want to build yourself.

Reach for FAISS when you want a fast embedded engine, when your data fits comfortably in memory on one machine, when you are learning how vector search works, or when you are building the retrieval part of a system you are willing to wrap in your own service. For a student project, a research experiment, or a prototype RAG application with a few hundred thousand chunks, FAISS is an excellent fit and has no running costs.

A note for readers working for Gulf and Egyptian employers: because FAISS runs entirely inside your own process, no data leaves your machine or your cloud region when you search. That is genuinely useful when data-residency rules apply. The embedding step is different. If you create embeddings by calling a hosted API, that step sends text to the provider, so check where the provider processes it.

Learning FAISS first makes the databases easier The words in the vector-database guides (index, nprobe, HNSW, quantization, recall) are the words you will meet here. Learn them once in FAISS and you can read any vector database's documentation.
Try it
  1. List a project idea of your own. Decide whether it needs metadata filters, network access, or live updates.
  2. If the answer is "no" to all three, write one sentence on why FAISS alone is enough.
  3. If "yes" to any, name which of the databases linked above you would look at first.

The mental model: four nouns

Almost everything in FAISS comes down to four ideas. If you hold these in your head, the API stops looking like a list of strange class names.

Vector. A row of float32 numbers. FAISS expects a whole collection of them at once as a two-dimensional NumPy array with shape (n, d): n rows, one per item, and d columns, the number of dimensions. The type must be float32, not NumPy's default float64. The value d is called the dimensionality and is fixed when you create an index. Every vector you add or search with must have exactly that many columns.

Index. The central object. An index holds your stored vectors (sometimes in compressed form) and knows how to search them. You create one by choosing a class, such as IndexFlatL2, give it d, and then call methods on it. Different index classes make different trade-offs between speed, memory, and accuracy, but they share almost the same methods: add, search, and a few others. This is why learning one index teaches you most of the others.

Distance (or similarity). How "close" is measured. The two you need now are L2 distance, the straight-line Euclidean distance, and inner product, a similarity where bigger means more alike. FAISS reports L2 as a squared distance, with no square root taken. We will spend a full section on this because it is where beginners meet their first confusing result.

Id. A 64-bit integer that names a stored vector. By default, ids are simply the order in which you added the vectors: the first is 0, the second is 1, and so on. A search returns ids, not vectors and not text.

Put together, the lifecycle of an index looks like this:

CREATEpick a class and d
→
TRAINonly some indexes
→
ADDstore vectors
→
SEARCHget distances and ids
→
SAVEwrite to a file

The train step needs explaining. Simple indexes know everything they need from birth. Cleverer ones, such as IVF and product quantization, must first learn something from a sample of your data, for example where the clusters are. That learning is called training, and you call index.train(sample) once before adding vectors. Every index has an attribute is_trained that tells you whether it is ready. Training happens only once: FAISS does not support retraining a populated index, so if your data changes character, you build a new index.

Two small attributes will become your most-used debugging tools: index.d, the dimension, and index.ntotal, the number of vectors stored. When something looks wrong, print these two before anything else.

Try it
  1. Without running code, say what shape a NumPy array should have to represent 5,000 sentences embedded with a 384-dimension model.
  2. Say what dtype it must have.
  3. Say what you would expect index.ntotal to print after adding it to an empty index.

Installing FAISS and checking the setup

For learning, you want the CPU package. It runs on Linux, macOS, and Windows and needs no special hardware. Create a clean virtual environment first so that you do not pollute your system Python or pick up an older FAISS by accident.

BASH
python -m venv .venv
source .venv/bin/activate          # on Windows: .venv\Scripts\activate
python -m pip install --upgrade pip
python -m pip install "faiss-cpu==1.15.1" numpy

Pinning the version with == is a good habit for any library you rely on, because FAISS releases every one to three months and some releases change behaviour.

There is a point worth stating plainly because older tutorials get it wrong. Until recently, the faiss-cpu package on PyPI was built by volunteers outside the project, and many guides told readers that the official route was conda only. That has changed. Since version 1.14.2 the FAISS project publishes official pip wheels, and the PyPI packages faiss-cpu, faiss-gpu and faiss-gpu-cuvs are maintained from the main repository. Pip is therefore a perfectly good way to install it. Conda is still described by the project's INSTALL.md as the supported route, and works like this:

BASH
conda install -c pytorch -c conda-forge faiss-cpu=1.15.1

On macOS with conda, note that the package supports Apple Silicon (arm64). Pip wheels also exist for Intel Macs on recent macOS versions.

A few platform facts that save hours:

  • There is no official GPU package for Windows or macOS. The GPU packages are for Linux with an NVIDIA card. On Windows, WSL2 with a Linux environment is the practical route if you want GPU later, but you do not need a GPU for anything in this guide.
  • Do not install faiss-cpu and faiss-gpu in the same environment. Both provide a Python module called faiss, and they will fight.
  • Install into the interpreter you actually run. A very common first error is ModuleNotFoundError: No module named 'faiss', which almost always means you installed with one Python and ran with another.

Now verify the install. Save this as check_faiss.py:

check_faiss.py
import faiss
import numpy as np

print("faiss version:", faiss.__version__)
print("compile options:", faiss.get_compile_options())
print("gpus visible:", faiss.get_num_gpus())

d = 64
xb = np.random.rand(1000, d).astype("float32")

index = faiss.IndexFlatL2(d)
index.add(xb)

D, I = index.search(xb[:5], 4)
assert (I[:, 0] == np.arange(5)).all()
print("ntotal:", index.ntotal)
print("nearest ids for the first query:", I[0])

Run it with python check_faiss.py. You should see the version 1.15.1, a line of compile options that mentions the CPU instruction set in use, gpus visible: 0 on a CPU package, and ntotal: 1000. The assertion checks something intuitive and useful: if you search the database with vectors that are already in it, the nearest neighbour of each one must be itself, at distance zero. That "search for yourself" check is a quick sanity test you will use again.

If the first line fails with ModuleNotFoundError, run python -m pip show faiss-cpu and compare the interpreter path to the one you ran. If it fails with an import error about NumPy, upgrade NumPy; the pip wheels need NumPy 1.25 or newer.

Do not use a package called faiss The PyPI name faiss is not this library. The official project names are faiss-cpu, faiss-gpu and faiss-gpu-cuvs. Third-party GPU wheels with names like faiss-gpu-cu12 are community builds, not official ones. Install only what the project documents.
Try it
  1. Create a fresh virtual environment and install FAISS with the pinned version.
  2. Run check_faiss.py and confirm every line prints.
  3. Change d to 65 and rerun. It still works, because d can be any positive number for a flat index. Keep this in mind for the product-quantization section later.

Your first working index, step by step

We will build a small but real index and walk through what each line does. We start with random vectors, because that removes everything except FAISS itself from the picture. Later in the guide we replace them with real embeddings.

Step 1: make the data. FAISS wants a NumPy array of shape (n, d) and dtype float32.

first_index.py
import faiss
import numpy as np

d = 128                      # dimension of every vector
n_db = 100_000               # number of stored vectors
n_q = 10                     # number of queries

rng = np.random.default_rng(42)
xb = rng.random((n_db, d), dtype="float32")   # the database
xq = rng.random((n_q, d), dtype="float32")    # the queries

Using a seeded random generator makes your results repeatable, which matters when you compare two runs. Notice that we asked for float32 directly. If you create data the default way, np.random.random(...) gives float64, and you must convert with .astype("float32"). The high-level Python functions convert some inputs for you, but relying on that hides real bugs, so make float32 a reflex.

Step 2: create the index. IndexFlatL2 is the simplest index. It stores every vector exactly as given and, at search time, compares the query with every one of them.

first_index.py
index = faiss.IndexFlatL2(d)
print(index.is_trained)      # True: a flat index needs no training
print(index.ntotal)          # 0: nothing stored yet

The name decodes as follows. Flat means "stored as is, no compression, no shortcuts". L2 means the distance is Euclidean. There is a sibling, IndexFlatIP, that uses inner product, which we cover in a later section.

Step 3: add the vectors. One call, and the ids are assigned for you.

first_index.py
index.add(xb)
print(index.ntotal)          # 100000

The first vector in xb gets id 0, the second id 1, and so on. Adding in several calls continues the numbering, so if you add 1,000 vectors and then 500 more, the second batch gets ids 1,000 to 1,499. FAISS does not copy these ids anywhere you can edit; they are simply positions.

Step 4: search. Ask for the 5 nearest neighbours of each of the 10 queries.

first_index.py
k = 5
D, I = index.search(xq, k)
print(D.shape, I.shape)      # (10, 5) (10, 5)
print(I[0])                  # ids of the 5 nearest vectors to query 0
print(D[0])                  # their squared L2 distances, smallest first

search returns two arrays, and learning to read them is the core skill of this whole guide.

  • I holds ids, with shape (n_q, k). Row i lists the ids of the k nearest stored vectors to query i, best first. Its dtype is int64.
  • D holds distances, with the same shape. D[i][j] is the distance between query i and the vector whose id is I[i][j].

To go from an id back to the vector itself in a flat index, you can index your own array: xb[I[0][0]]. FAISS also offers index.reconstruct(i) for some index types, but because you usually keep the original data anyway, indexing your own array is simpler.

Step 5: sanity-check against yourself. Search with vectors you know are in the database.

first_index.py
D_self, I_self = index.search(xb[:3], 1)
print(I_self.ravel())        # [0 1 2]
print(D_self.ravel())        # [0. 0. 0.] (or extremely close to 0)

If the self-search does not return the right ids with distance near zero, something is wrong with your data pipeline, not with FAISS. Hold on to that diagnostic.

Step 6: time it. Try a larger query batch and measure.

first_index.py
import time

xq_big = rng.random((1_000, d), dtype="float32")
t0 = time.perf_counter()
D, I = index.search(xq_big, 10)
print(f"{time.perf_counter() - t0:.3f} s for 1000 queries")

You will see that a flat search over 100,000 vectors is fast on a normal laptop, because it uses optimised matrix operations over a whole batch. That is an important lesson about FAISS. It is built for batches. Searching 1,000 queries in one call is much faster per query than calling search a thousand times with one query each. When you build an application, collect queries into batches where you can.

A single query must still be two-dimensional index.search(q, 5) with q of shape (128,) fails with an assertion error about the dimension. Reshape it into one row: q.reshape(1, -1) gives shape (1, 128). This is the most common first error, and it is covered again in the errors section.
Try it
  1. Type in first_index.py and run it.
  2. Change k to 1, then to 50, and check that the shapes of D and I follow.
  3. Ask for k = 200000, which is more than the database holds. Look at what appears in I and D. You should see -1 ids, which is FAISS's way of saying "there was no result to put here".

Reading distances: L2, inner product, and cosine similarity

You now have a search that returns numbers, so the next question is what the numbers mean. This is the most conceptually tricky part for beginners, and a few minutes of care here prevents a lot of silent wrong answers.

L2 distance. IndexFlatL2 uses METRIC_L2. The value in D is the squared Euclidean distance: the sum of the squared differences between the two vectors' components, with no square root applied. Smaller means closer, and an exact match gives 0. FAISS skips the square root because it does not change the order of results and costs time. If you need the true distance, take np.sqrt(D) yourself.

Inner product. IndexFlatIP uses METRIC_INNER_PRODUCT. The value is the dot product of the query and the stored vector. Here bigger means more similar, and the results come back sorted from highest to lowest. This reversal trips people up: with L2 you read the smallest number first, with inner product the largest.

Cosine similarity. Many embedding models are designed to be compared by cosine similarity, which measures the angle between two vectors and ignores their length. FAISS has no cosine index. The standard recipe is simple and worth memorising: normalise every vector to length 1, then use inner product. For unit-length vectors, the inner product is the cosine similarity.

cosine.py
import faiss
import numpy as np

d = 128
rng = np.random.default_rng(0)
xb = rng.random((10_000, d), dtype="float32")
xq = rng.random((5, d), dtype="float32")

faiss.normalize_L2(xb)       # in place! xb is modified
faiss.normalize_L2(xq)       # normalise queries the same way

index = faiss.IndexFlatIP(d)
index.add(xb)
D, I = index.search(xq, 5)
print(D[0])                  # cosine similarities, in [-1, 1], largest first

Three details in that snippet matter.

First, faiss.normalize_L2 works in place. It changes the array you pass in and returns nothing, so there is no xb = faiss.normalize_L2(xb). It also requires a float32 array, which is another reason to get the dtype right from the start.

Second, you must normalise both the stored vectors and the queries, with the same function. If you normalise only one side, your "cosine" scores are meaningless, and the results will look plausible but be wrong. This is the single most common silent bug in beginner FAISS code.

Third, some embedding models already return unit-length vectors. Normalising a vector that is already unit length does no harm, so when in doubt, normalise.

For unit-length vectors, L2 and cosine rank results identically, because the squared L2 distance equals 2 - 2 x cosine. So an L2 index on normalised vectors gives the same ordering as an inner-product index. Use whichever you find easier to explain.

How do you decide which metric to use? Read the model card of your embedding model. If it says to use cosine similarity or dot product with normalised embeddings, normalise and use IndexFlatIP. If it says Euclidean distance, use IndexFlatL2. Using the metric a model was trained for gives the best results.

Check your normalisation After normalising, np.linalg.norm(xb, axis=1) should be all approximately 1. Print its minimum and maximum once. It takes ten seconds and catches the "I normalised the wrong array" mistake.
Try it
  1. Build the same data twice: once with IndexFlatL2 on normalised vectors, once with IndexFlatIP on the same normalised vectors.
  2. Search with the same queries and confirm the returned ids match.
  3. Take one pair of results and verify that the squared L2 distance equals 2 - 2 * similarity.

Ids, custom ids, and keeping your own data

By default, FAISS ids are just positions: 0, 1, 2 and so on. That is fine if your documents live in a Python list or a table in the same order. You search, get ids, and index into your list.

lookup.py
documents = ["reset your password", "update billing details", "export your data"]
# ... suppose these were embedded, in this order, and added to the index ...
# D, I = index.search(query_vector, 2)
# for doc_id in I[0]:
#     print(documents[doc_id])

That pattern is the entire "storage" story for beginners: a list or a database table whose row number equals the FAISS id. If you ever delete or reorder the list without rebuilding the index, the ids point at the wrong things. That is a silent, nasty bug, so keep the list and the index together, and save them together.

Sometimes you want your own ids, for example database primary keys such as 10042 and 10043. A flat index does not accept them directly. You wrap the index in an IndexIDMap:

custom_ids.py
import faiss
import numpy as np

d = 64
rng = np.random.default_rng(1)
xb = rng.random((5, d), dtype="float32")
ids = np.array([101, 102, 103, 104, 105], dtype="int64")

index = faiss.IndexIDMap(faiss.IndexFlatL2(d))
index.add_with_ids(xb, ids)

D, I = index.search(xb[:1], 3)
print(I)        # ids drawn from 101..105, never 0..4

Pay attention to the details. Ids must be a NumPy array of dtype int64. FAISS has no string ids and never will, so a product code such as "SKU-9912" must be mapped to an integer through a dictionary you maintain. You wrap the index before adding any vectors; wrapping a non-empty index fails with an error that says the index must be empty on input. And if you call plain add on an IndexIDMap, it raises an error, because it needs ids; use add_with_ids.

If you try add_with_ids directly on IndexFlatL2 you will get add_with_ids not implemented for this type of index. Flat and HNSW indexes only know sequential ids. Wrapping them with IndexIDMap fixes that, and the index factory string for the same thing is "IDMap,Flat".

Finally, remember that FAISS stores no metadata. If a search result is "id 4017" you need your own table to learn that id 4017 is a document called "Refund policy, v3", written in Arabic, tagged "billing". A simple Python dictionary, a CSV file or a database table all work for a learning project.

Try it
  1. Create ten short strings in a list and embed them as random vectors (it is fine to fake the embedding for now).
  2. Add them with custom ids that start at 5000, using IndexIDMap.
  3. Search with one of the stored vectors, map the returned ids back to your strings, and check you get the same string first.

Saving and loading an index

An index in memory disappears when the process ends. Saving takes two functions.

persist.py
import faiss

faiss.write_index(index, "vectors.index")
index2 = faiss.read_index("vectors.index")
print(index2.ntotal, index2.d)

write_index writes the whole index to a single file, including anything it learned during training. read_index gives you back an equivalent index that you can search immediately. For a flat index of 100,000 vectors of 128 numbers, the file is about 51 MB, because every vector is stored as four bytes per number, plus a small header. You can check that with ls -lh vectors.index.

A handful of practical points about saving:

  • The directory must already exist, or you get could not open ... for writing. A wrong path on loading gives could not open ... for reading: No such file or directory.
  • Save your id-to-data table beside the index, in the same folder and version. A restored index is useless if the list that explains its ids is missing or has changed.
  • A newer FAISS can read files written by older versions, but not necessarily the other way round. If you write an index with a recent FAISS and try to read it with an old one, you may see Index type ... not recognized. Pin the same version in every environment that touches the file.
  • GPU indexes cannot be written directly. You convert to a CPU index first. This does not matter on the CPU package, but remember it for later.

There is also a security point that is easy to miss and important to learn on day one. The FAISS project's own documentation warns that no attempt is made to verify that a loaded index file is valid, and that a faulty or malicious file could cause out-of-memory errors or, if crafted expertly, code execution. Recent releases hardened the reading code considerably, but the rule remains: load only index files you created yourself or received from a source you trust, and never load one a stranger sent you. The same applies to the files that some wrapper libraries save next to the index, which can contain Python pickles.

An index file is not just data Treat .index files like executable code from the point of view of trust. Download them only from places you control, and keep a checksum (for example SHA-256) of the file you built, so you can tell if it has changed.
Try it
  1. Save the index from the first project, then restart Python entirely.
  2. Load it, run the same queries as before, and confirm you get identical ids.
  3. Delete the file, run the load again, and read the exact error text so that you recognise it later.

Making search faster: the IVF index

Flat search is exact, and for up to a few hundred thousand vectors it is often all you need. Past that, it becomes the bottleneck, because every query compares against every vector. The first and most widely used shortcut is the inverted file index, written IVF.

The idea is easy to picture. Imagine a library where books are sorted into rooms by topic. To find a book about gardening, you do not walk through every room; you go to the gardening room and perhaps one or two next to it. IVF does this with vectors. At training time it runs k-means clustering on a sample of your data to find nlist cluster centres (called centroids), and each region of space closest to one centroid is a cell. When you add vectors, each is filed in the cell of its nearest centroid. At search time, FAISS finds the nprobe cells closest to the query and only compares against vectors in those cells.

TRAINk-means finds nlist cells
→
ADDeach vector goes in a cell
→
SEARCHlook in nprobe nearest cells

This produces the central trade-off of approximate search. A smaller nprobe is faster but can miss a true neighbour that sat in a cell you did not visit. A larger nprobe is slower and more accurate. When nprobe equals nlist, you visit every cell and the search is exhaustive again, which gives the same answers as the flat index.

ivf.py
import faiss
import numpy as np

d = 128
rng = np.random.default_rng(7)
xb = rng.random((100_000, d), dtype="float32")
xq = rng.random((10, d), dtype="float32")

nlist = 1024
quantizer = faiss.IndexFlatL2(d)               # decides which cell a vector belongs to
ivf = faiss.IndexIVFFlat(quantizer, d, nlist)  # metric defaults to L2

print(ivf.is_trained)     # False: IVF must learn its cells first
ivf.train(xb)             # k-means over the data
print(ivf.is_trained)     # True

ivf.add(xb)
ivf.nprobe = 16           # the default is 1, which is too low for good recall
D, I = ivf.search(xq, 5)

Read this carefully, because three points confuse beginners every time.

The quantizer is a small separate index used only to find the nearest centroid. It is usually a flat index. It is an unusual name for it, but the role is simple: it assigns vectors to cells. Keep a reference to it alive as long as the IVF index is in use.

Training is mandatory. If you skip ivf.train(...) and call add you get Error: 'is_trained' failed. The training sample should resemble the data you will index. Here we trained on all of xb, but with larger data you train on a random subset. As a rule of thumb from the FAISS documentation, use somewhere between 30 and 256 training vectors per cell, and expect FAISS to print a warning if you give fewer than 39 per cell. For nlist = 1024 that means at least about 40,000 training vectors; the warning looks like WARNING clustering 1000 points to 1024 centroids: please provide at least 39936 training points.

nprobe defaults to 1, and that is the cause of most "my IVF index returns bad results" complaints. With one cell visited, you only see vectors in the single closest cell. Raise it to 8, 16, or 32 and compare. A good starting rule for the number of cells is between 4 and 16 times the square root of the number of vectors, so for 100,000 vectors, roughly 1,300 to 5,000. The official guidelines give this for collections up to a million vectors.

To measure how good an approximate index is, you need a ground truth: the answers from an exact flat search. The measure is recall: of the true top-k neighbours, what fraction did the approximate index also return?

recall.py
flat = faiss.IndexFlatL2(d)
flat.add(xb)
_, I_true = flat.search(xq, 5)

for nprobe in (1, 4, 16, 64):
    ivf.nprobe = nprobe
    _, I_ivf = ivf.search(xq, 5)
    hits = [len(set(a) & set(b)) for a, b in zip(I_true, I_ivf)]
    print(nprobe, sum(hits) / (5 * len(xq)))

You will see recall climb as nprobe grows. One honest caveat: random vectors have no cluster structure, so IVF performs worse on them than on real embeddings, and recall will be lower than you would see on real text or image data. Do not draw conclusions about real-world accuracy from random numbers; use them only to learn the mechanics.

Always keep a flat index for testing Whatever approximate index you build, keep a small flat index of the same data in your test script. It costs a little memory and gives you the exact answers against which every other setting is judged.
Try it
  1. Build the IVF index above and run the recall loop. Write down the recall for each nprobe.
  2. Time the searches for each nprobe and compare with the flat index. Note where the speed gain stops being worth the lost recall.
  3. Set nlist to 4096 with only 5,000 training vectors, and read the warning that FAISS prints.

Making search faster another way: the HNSW graph

IVF is one family. The other popular family is HNSW, which stands for hierarchical navigable small world. Instead of sorting vectors into rooms, it connects each vector to a handful of its neighbours, building a layered graph. To search, FAISS starts at an entry point and repeatedly hops to whichever neighbour is closer to the query, like finding a street by asking people on each corner whether they know the way. The upper layers are sparse and allow big jumps; the bottom layer is dense and allows precise finishing.

hnsw.py
import faiss
import numpy as np

d = 128
rng = np.random.default_rng(3)
xb = rng.random((100_000, d), dtype="float32")
xq = rng.random((10, d), dtype="float32")

index = faiss.IndexHNSWFlat(d, 32)       # M = 32 links per vector
index.hnsw.efConstruction = 64           # set BEFORE adding (default 40)
index.add(xb)                            # no training needed
index.hnsw.efSearch = 64                 # default is 16
D, I = index.search(xq, 5)

There are three settings to understand. M is the number of connections per vector: more connections give better accuracy and use more memory. efConstruction controls how carefully the graph is built, and it only matters while adding, so set it before add. efSearch controls how wide the search is at query time, and it plays the same role as nprobe: raise it for better recall and slower queries. The defaults in FAISS are efConstruction = 40 and efSearch = 16.

HNSW has a different personality from IVF, and knowing it helps you choose:

  • No training. You can add vectors straight away, which makes it pleasant to use.
  • Fast and accurate at the cost of memory. Besides the vectors themselves, each one carries link storage of roughly 256 bytes for M = 32, so HNSW suits cases where RAM is plentiful.
  • No deletion. remove_ids is not supported on HNSW indexes. If you must remove items, rebuild the index or use another type.
  • Sequential ids only. To use your own ids, wrap it with IndexIDMap, which brings the IDMap rules from earlier.
Try it
  1. Build the HNSW index and compare its recall against the flat index using the same loop as for IVF, but varying efSearch over 8, 16, 64, 256.
  2. Time add for HNSW and for IVF (including training). Which is cheaper to build?
  3. Try index.remove_ids(...) on the HNSW index and read the error.

The index factory: building indexes from a string

Writing the constructor classes by hand gets awkward as indexes become more elaborate. FAISS provides a compact text language called the index factory. You describe the index in a short string and get back the right object.

factory.py
import faiss

d = 128
a = faiss.index_factory(d, "Flat")                 # same as IndexFlatL2(d)
b = faiss.index_factory(d, "IVF1024,Flat")         # same as IndexIVFFlat with nlist=1024
c = faiss.index_factory(d, "HNSW32")               # same as IndexHNSWFlat(d, 32)
e = faiss.index_factory(d, "IVF1024,Flat", faiss.METRIC_INNER_PRODUCT)
f = faiss.index_factory(d, "IDMap,Flat")           # flat index with custom ids

Read the strings left to right as a pipeline: optional pre-processing, then the structure, then how vectors are stored. "IVF1024,Flat" means "IVF with 1024 cells, storing full vectors". "HNSW32" means "HNSW with 32 links". The third argument sets the metric and defaults to L2, so remember to pass faiss.METRIC_INNER_PRODUCT when working with cosine similarity.

A bad string gives could not parse index string .... Some strings carry hidden requirements: for example, "IVF4096,PQ16" needs d to be divisible by 16. You will meet compression codes like PQ and SQ8 in the next level of this series; for now you only need the Flat, IVF...,Flat, HNSW... and IDMap,... forms.

A useful trick that avoids a famous trap: with an index created by index_factory, you may not know its concrete class, so setting attributes such as index.nprobe = 16 can silently do nothing. The reliable way to set search settings on any index is ParameterSpace:

params.py
ps = faiss.ParameterSpace()
ps.set_index_parameters(b, "nprobe=16")      # works on factory-built or wrapped indexes

When the index is wrapped (for instance in an IndexIDMap), assigning a field directly on the wrapper creates a harmless Python attribute and changes nothing in the C++ object, so searches stay at the default. ParameterSpace reaches through the wrapper and sets the real value.

Silent no-ops are the worst kind of bug If you set nprobe or efSearch and the recall does not change at all, suspect that you set it on a wrapper. Use ParameterSpace and measure again.
Try it
  1. Build "IDMap,Flat" with the factory and add vectors with custom ids.
  2. Build "IVF256,Flat", train it, and set nprobe with ParameterSpace.
  3. Feed the factory a wrong string, such as "IVF,Flat", and read the message.

Choosing an index when you are a beginner

With three families on the table, you need a practical way to choose. The FAISS project publishes "Guidelines to choose an index", and the beginner-relevant parts boil down to a short decision list. Think of it as a ladder: start at the top and only step down when you have a measured reason.

  1. Up to roughly a few hundred thousand vectors, or when you need exact results: Flat. It needs no tuning, no training and has no accuracy loss. If a flat search is fast enough for your application, stop here. The official guidance says that for a few thousand searches, or anytime exact answers are required, flat is the answer.
  2. Plenty of memory and no deletions: HNSW. HNSW32 is a strong default. It is fast and accurate and needs no training.
  3. Larger data, or memory is somewhat tight: IVFK,Flat. Choose K as 4 to 16 times the square root of your number of vectors, train it on 30 to 256 vectors per cell, and raise nprobe until recall is acceptable.

Beyond roughly a million vectors, the guidelines recommend IVF with an HNSW coarse quantizer, such as IVF65536_HNSW32 for one to ten million vectors, and compression to cut memory. Those are mid-level and senior topics, and the later parts of this series pick them up.

It helps to do the memory arithmetic before you start, because RAM is usually the limit that decides things. A flat index stores 4 x d bytes per vector. For 1 million vectors of 768 numbers, that is about 3 GB. At 100 million vectors it is about 307 GB, which does not fit on a laptop and will be the reason to learn compression later. HNSW adds around 256 bytes per vector for links, and IVF adds 8 bytes per vector for the id.

Try it
  1. Compute, on paper, the memory of a flat index for 2 million vectors of 384 numbers. Check with index.ntotal * d * 4 bytes.
  2. For a dataset of your own choosing, decide which rung of the ladder fits and say why in two sentences.
  3. Build that index on random data of the same size and record its build time and search time.

Useful helpers: brute-force search and k-means

Two helpers deserve a mention, because they show that FAISS is more than indexes. The first is a brute-force search with no index object at all, useful for quick comparisons and for checking an index.

helpers.py
import faiss
import numpy as np

d = 32
rng = np.random.default_rng(5)
xb = rng.random((20_000, d), dtype="float32")
xq = rng.random((4, d), dtype="float32")

D, I = faiss.knn(xq, xb, 5, metric=faiss.METRIC_L2)   # exact top 5, no index to build
dist = faiss.pairwise_distances(xq, xb)               # a full (4, 20000) distance matrix

faiss.knn is handy in notebooks and tests: it returns exactly what a flat index would. pairwise_distances gives every distance between two sets and is only suitable for modest sizes, because the result has one entry per pair.

The second helper is k-means clustering, the same algorithm IVF uses internally, offered directly:

kmeans.py
km = faiss.Kmeans(d, 10, niter=20, verbose=True)   # 10 clusters, 20 iterations
km.train(xb)
print(km.centroids.shape)                          # (10, 32)
D, labels = km.index.search(xb, 1)                 # nearest centroid per vector

After training, km.centroids holds the cluster centres, and searching km.index with your data gives each vector's cluster. This is a fast way to group embeddings, for example to see themes in a pile of customer messages. The default number of iterations is 25 if you do not pass niter. As with IVF, if you ask for many clusters from too few points, FAISS warns you that you should provide more.

Try it
  1. Compare faiss.knn with a flat index on the same data and confirm the two return identical ids.
  2. Cluster 20,000 random vectors into 5 groups and count how many fall in each. Run it twice with different seeds and compare.

Configuration you will actually touch: threads and data types

FAISS has few configuration files, because it is a library, but a few settings matter even for beginners.

Threads. FAISS uses OpenMP to run a batch of queries in parallel across CPU cores. You normally leave it alone. When you need to limit it, for instance on a shared machine or inside a container, there are two ways:

BASH
export OMP_NUM_THREADS=4
python first_index.py

or in Python, faiss.omp_set_num_threads(4). A subtlety makes the second form surprising: the setting applies to the current thread only. If you call it in one thread and run searches in another, it appears to have no effect. The environment variable has no such problem, so prefer it in scripts and services. To see the current value, call faiss.omp_get_max_threads().

If searches are mysteriously slow on a machine that uses OpenBLAS, the cause can be two layers of threading fighting each other. Setting OMP_WAIT_POLICY=PASSIVE is the documented remedy.

Set threads before you benchmark When comparing two index types, fix OMP_NUM_THREADS to the same value for both. Otherwise a difference in timing may come from the number of threads and not from the index.
Try it
  1. Run the 1,000-query timing from earlier with OMP_NUM_THREADS=1 and again with OMP_NUM_THREADS=4. Compare.
  2. Print faiss.omp_get_max_threads() before and after setting the variable.

Common errors and how to read them

FAISS is written in C++, and its exceptions arrive in Python as RuntimeError, with the message in the form Error in <function> at <file>:<line>: <message>. The useful part is the end. Some failures are harder: a hard assertion prints Faiss assertion '<condition>' failed and kills the whole process, with no chance to catch it. That is a crash, not an exception. Use the table to turn messages into fixes.

Symptom or message Real cause Fix
AssertionError on d == self.d The array's second dimension is not the index's d Print x.shape and index.d; reshape a lone query with q.reshape(1, -1)
Error: 'is_trained' failed add or search on an IVF or similar index that has not been trained Call index.train(sample) first and check index.is_trained
WARNING clustering ... please provide at least ... training points Fewer than 39 training points per cell More training data, or a smaller nlist
add_with_ids not implemented for this type of index Flat and HNSW store sequential ids only Wrap in faiss.IndexIDMap(...) or use the "IDMap,Flat" factory string
remove_ids not implemented for this type of index HNSW cannot delete Rebuild, or use a Flat, IVF or IDMap index
could not parse index string Bad factory string Compare it to the forms in this guide; check that d fits any PQ or SQ part
could not open ... for reading Wrong path to read_index Check the path and the working directory
Index type ... not recognized Corrupt file, binary index read with read_index, or a file from a newer FAISS Use the matching reader or version
ModuleNotFoundError: No module named 'faiss' Installed in a different interpreter python -m pip install faiss-cpu with the interpreter you run
TypeError ... argument ... of type 'faiss::idx_t' A NumPy integer where Python expects an int Wrap with int(...)
-1 ids and huge distances in results k larger than ntotal, or the search visited too few candidates Lower k; raise nprobe or efSearch
Poor recall from an IVF index nprobe still at default 1 Raise nprobe and re-measure
Setting nprobe changes nothing Set on a wrapper object Use ParameterSpace().set_index_parameters(...)
Error: 'index_->ntotal == 0' failed Wrapping a non-empty index in IndexIDMap Wrap first, add afterwards
AttributeError: module 'faiss' has no attribute 'StandardGpuResources' You are on the CPU package Expected on CPU; ignore GPU code, or install a GPU package on Linux
Reading errors top to bottom is the wrong order With C++ exceptions, the first line shows where the code was, but the last words after the final colon tell you what is wrong. Read from the right.
Try it
  1. Cause three errors on purpose: an untrained IVF add, a wrong-dimension query, and add_with_ids on a flat index.
  2. For each, copy the message into a notes file with its fix written beside it. You now have your own troubleshooting table.

Putting it all together: a small semantic search project

Time to combine everything into one useful program: a tiny semantic search engine over a handful of support questions. It uses sentence-transformers to produce real embeddings, which is a separate library you install alongside FAISS. Real embeddings are what make the results meaningful, and they also let you see why normalisation matters.

First install the extra library and set up a folder:

BASH
python -m pip install sentence-transformers
mkdir faiss_search && cd faiss_search

Now the program. It embeds documents, builds a cosine-similarity index, saves it with its document list, and answers queries.

search_app.py
import json
import faiss
import numpy as np
from sentence_transformers import SentenceTransformer

docs = [
    "How do I reset my password?",
    "I forgot my login credentials",
    "How can I update my billing address?",
    "Where do I download my invoices?",
    "How do I cancel my subscription?",
    "The app crashes when I open it",
    "How do I export all my data?",
    "Can I change my account email?",
]

model = SentenceTransformer("all-MiniLM-L6-v2")      # produces 384-number vectors
xb = model.encode(docs).astype("float32")             # shape (8, 384)
faiss.normalize_L2(xb)                                # unit length, in place

d = xb.shape[1]
index = faiss.IndexFlatIP(d)                          # inner product on unit vectors = cosine
index.add(xb)

faiss.write_index(index, "support.index")
with open("support_docs.json", "w") as f:
    json.dump(docs, f)

def ask(question, k=3):
    q = model.encode([question]).astype("float32")    # note the list: shape (1, 384)
    faiss.normalize_L2(q)
    D, I = index.search(q, k)
    for score, i in zip(D[0], I[0]):
        print(f"{score:.3f}  {docs[i]}")

ask("I can't sign in to my account")
ask("where can I find my receipts")

Walk through what happens. The model turns each sentence into a 384-number vector (that dimension is a property of this particular model, so d is read from the data instead of being typed in). We convert to float32 and normalise in place. The index uses inner product, which on unit vectors equals cosine similarity. We save two files together: the FAISS index and a JSON list of the documents whose positions are the ids.

The ask function is where the lessons pay off. A single question is encoded as a list of one string, which gives shape (1, 384) and avoids the dimension error. We normalise it the same way as the stored vectors. Then search returns scores in D and ids in I, and we use the ids to look up the original text from docs.

When you run it, "I can't sign in to my account" should rank the password and login questions first, even though it shares hardly any words with them. That is the point of embeddings. "Where can I find my receipts" should surface the invoice question. If the first run downloads a model, it needs internet once and then caches it. Remember that embedding a text with this local model happens entirely on your machine, which keeps your documents in place; a hosted embedding API would send them to the provider.

To reuse the saved index later, in a fresh process, load both files and you never need to embed the documents again:

reload.py
import json
import faiss

index = faiss.read_index("support.index")
with open("support_docs.json") as f:
    docs = json.load(f)
print(index.ntotal, len(docs))      # these two numbers must be equal

That final comparison is a good habit: the number of vectors and the number of documents must match, and a mismatch means the two files are out of sync.

To grow this into something realistic, change three things. Embed in batches rather than one at a time. Switch the index to HNSW32 or IVF... once you have enough data to need speed, measuring recall against flat. And add a ParameterSpace setting to tune nprobe or efSearch. If you later want to feed these retrieved documents to a language model for question answering, that is a RAG pipeline, and frameworks such as LangChain and LlamaIndex provide FAISS wrappers. Those wrappers belong to those projects, so check their documentation for current signatures and be careful with how they save files, since some use pickling.

Try it
  1. Run search_app.py and add five more documents of your own, in a language you work in. The model named above is trained mainly on English; try a multilingual model for Arabic text and compare.
  2. Write a query that is far from every document and look at the scores. Decide what threshold of similarity you would use to say "no good match".
  3. Reload the saved files in a new process and ask three questions without re-embedding the documents.

What you can now do, and what comes next

You can now do real work with FAISS. You know that it is an embedded library for nearest-neighbour search over float32 vectors, and not a database. You can install an official, pinned version, verify it, build a flat index, add vectors, search them, and read the two arrays that come back. You understand the difference between squared L2, inner product, and cosine, and you know the normalise-both-sides recipe. You can use custom ids through IndexIDMap, keep a side table for your real data, and save and load an index while treating the file as something to trust carefully. You can speed up search with IVF and HNSW, measure their recall against a flat baseline, use the index factory, and set search parameters safely. You can read the common error messages in the table above.

What comes next, in the later parts of this series:

  • Mid-level covers compression (product quantization, scalar quantization, FastScan), re-ranking with refinement, per-query search parameters that are safe across threads, filtering by id with selectors, deletes and updates, and tuning with ParameterSpace.
  • Senior covers running FAISS as a service, GPUs and cuVS, indexes that do not fit in RAM, sharding, upgrades, security, multi-tenancy, and when to move to a vector database.

Beyond this series, look at Qdrant, Milvus, Weaviate, Chroma, Pinecone and pgvector once you need persistence, metadata filters, or a network API. For evaluation of the retrieval step in a RAG system, look at RAGAS. Having built a FAISS index by hand, you will understand precisely what each of those systems is doing for you.

Try it
  1. Write your own one-page cheat sheet with the five calls you used most: IndexFlatL2, add, search, write_index, read_index.
  2. Add three "never forget" lines of your own: float32, normalise both sides, nprobe defaults to 1.

Sources