تخطَّ إلى المحتوى
العودة إلى أدلة الدارسين
LlamaIndexLLMsFrameworks & agents3 مستويات91 قسمًايغطّي LlamaIndex 0.14دليل بالإنجليزية

The Complete LlamaIndex Guide

Connect LLMs to your own data: ingestion, indexing, retrieval and agentic RAG. Taught at three levels — Beginner, Mid-level and Senior — each with an in-depth guide, interview prep, and practical tips.

التوثيق الرسمي مسودّة بالذكاء الاصطناعي · مراجعة المجتمع جاريةساعدنا في مراجعته
13sections
21examples

This is part one of three. It covers everything you need to build a working question-answering application over your own documents, and a first agent that can call your own Python functions. By the end you can install LlamaIndex correctly, load a folder of files, build an index, ask questions and get answers grounded in those files, save the index so you never pay to rebuild it twice, swap OpenAI for a model running on your own laptop, and read the error messages the framework throws at you. Mid-level and Senior take the same ideas into production; nothing here is wasted.

Every section ends with a Try it task. Do them while you read. Retrieval is one of those subjects that feels obvious on the page and surprises you the moment you run it on real documents, so the sooner you see your own output the better.

0.14.25llama-index-core, 21 Sep 2026
Python 3.10–3.143.9 is no longer supported
1024 / 200default chunk size / overlap, in tokens
top_k = 2chunks retrieved by default

What LlamaIndex is, and the problem it solves

LlamaIndex is a Python framework for building applications that let a large language model answer questions about your data. The official description is "a framework for building LLM-powered applications over your data". That phrase is doing a lot of work, so unpack it.

A language model such as GPT-4o or Claude knows what was in its training data and nothing else. It has never seen your company's HR handbook, the 400 PDFs your compliance team keeps on a shared drive, your product's API documentation, or the Arabic-language policy circulars your bank published last quarter. Ask it a question about any of those and one of two things happens: it tells you it does not know, or — worse, and far more common — it produces a confident, fluent, completely invented answer. The industry word for that is hallucination, and for an internal tool it is fatal. A chatbot that is right 85% of the time and gives no hint about which 15% is wrong is a liability, not a product.

There are two ways to fix this. One is to train the model on your data, which is expensive, slow, has to be repeated every time a document changes, and still does not give you a citation you can show a reviewer. The other is to look the answer up at the moment the question arrives, and hand the model the relevant passages along with the question. The second approach is called RAG, retrieval-augmented generation. The LlamaIndex docs define it as a technique that "allows LLMs to answer questions about your private data by providing it to the LLM at query time, rather than training the LLM on your data", sending "only the relevant parts along with your query".

That last clause is the engineering problem in a nutshell. You cannot paste 400 PDFs into a prompt; models have a context window, and even when the window is large, cost and latency scale with every token you send, and accuracy degrades when the useful sentence is buried in a hundred pages of noise. So you need machinery that splits documents into pieces, works out which pieces are relevant to a question, and sends only those. LlamaIndex is that machinery, plus a consistent interface over the dozens of model providers, vector databases and file formats you might plug into it.

YOUR FILESpdf, md, docx
→
CHUNKSnodes
→
INDEXvectors
→
RETRIEVEtop matches
→
ANSWERLLM + sources

What came before, and why this shape won

Before frameworks like this existed, every team built the same pipeline by hand. You wrote a loop over a directory, picked a PDF library, invented a chunking rule, called an embedding API, pushed the vectors into whatever database you had, wrote a similarity query, and glued a prompt template on the end. It worked, and it took two weeks, and then you discovered your chunking cut tables in half, and the rewrite took another week.

LlamaIndex, which started life in 2022 as "GPT Index", packaged that pipeline. Its sibling frameworks — most obviously LangChain — made similar bets. The practical difference in 2026 is one of centre of gravity: LangChain grew out of chaining prompts and generalised towards agents, while LlamaIndex grew out of indexing documents and generalised towards agents from the data side. Both arrive in roughly the same place now, and the choice between them is usually about which abstractions you find clearer rather than what is possible.

What you get by using a framework instead of hand-rolling is substitution. Every component is behind an interface, so changing your embedding model from OpenAI's hosted one to a local HuggingFace model is one line, and moving from an in-memory store to Qdrant or pgvector is two. That matters enormously for employers in the Gulf and in Egypt, where "the documents may not leave the country" is a routine constraint: the same application code can run against a hosted model for a prototype and against a self-hosted model in an on-premise or in-region deployment for production, because the retrieval layer does not care.

Try it
  1. Pick a real document set you know well — course notes, a policy PDF, a repository's docs folder.
  2. Write down three questions whose answers are definitely in there, and one question whose answer is definitely not.
  3. Ask a chatbot all four without giving it the documents. Note which answers are wrong and whether you could tell.
a concrete list of failures. You will use these same four questions to test your own application later in the guide, and the fourth one is the important one: a good RAG system should decline it.

The mental model: six nouns

Almost everything in LlamaIndex is one of six things. Learn these names properly now and the API documentation stops being a maze.

A Document is a container for one whole source: a file, a web page, a database row. It holds text, a dictionary of metadata, and an id (id_, also exposed as doc_id). When you load a folder of twenty PDFs you get Documents back, one per page or per file depending on the reader.

A Node is a chunk of a Document. The docs call it "a lightweight abstraction over text strings that keep track of metadata and relationships". The common subclass is TextNode; there are also ImageNode and IndexNode. A Node inherits its parent's metadata and keeps a pointer home in node.ref_doc_id. Nodes, not Documents, are what actually gets embedded, retrieved and shown to the model. This is the single most important distinction in the framework: you load Documents, but you retrieve Nodes.

A Reader, also called a data connector, turns a source into Documents. The built-in one for local files is SimpleDirectoryReader. Hundreds more — Notion, Google Drive, S3, SharePoint, web pages — ship as separate packages.

A node parser (equivalently a text splitter, and more generally a transformation) turns Documents into Nodes. The default is SentenceSplitter, which tries to break on sentence boundaries. Others include TokenTextSplitter, SemanticSplitterNodeParser, CodeSplitter and MarkdownElementNodeParser.

An Index is a data structure over Nodes that serves retrieval. VectorStoreIndex is the default and the only one you need at this level: it embeds each Node into a vector and finds matches by similarity. Others exist for other shapes of question — SummaryIndex, DocumentSummaryIndex, PropertyGraphIndex, KeywordTableIndex, TreeIndex.

A query engine is the end-to-end single-turn flow: take a question, retrieve Nodes, optionally post-process them, and synthesise an answer. You get one with index.as_query_engine() and call it with .query(). Its stateful multi-turn cousin is the chat engine, index.as_chat_engine().

Three more terms complete the picture. A retriever is just the retrieval half of a query engine, returning NodeWithScore objects, obtained with index.as_retriever(). A node postprocessor filters or re-ranks retrieved Nodes before synthesis, for example SimilarityPostprocessor(similarity_cutoff=0.75). A response synthesizer turns Nodes plus the question into prose; its default mode for query engines is compact, with refine and tree_summarize as the main alternatives.

The second half of the framework

Everything above is the data layer, which lives in llama_index.core. There is a second, independent half: Workflows, shipped as llama-index-workflows and re-exported as llama_index.core.workflow. A Workflow is "an event-driven, step-based way to control the execution flow of an application". You will not write one by hand in this guide, but you need to know it exists for one reason: every agent in LlamaIndex is a Workflow underneath. That is why agents are asynchronous, and why their errors mention workflow concepts like timeouts and iteration limits.

An agent is, in the docs' words, "a specific system that uses an LLM, memory, and tools, to handle inputs from outside users". A tool is a callable plus metadata — a name, a description, a parameter schema — that the agent is allowed to invoke. The loop is simple: the agent sends the user's message plus the tool schemas to the model; the model either answers or asks for one or more tool calls; your code runs them and appends the results; repeat until the model answers. The classes you will meet are FunctionAgent (uses the provider's native tool-calling), ReActAgent (uses Thought/Action/Observation prompting, so it works with models that have no tool-calling API), CodeActAgent (writes and runs code) and AgentWorkflow (several agents handing off to each other).

Why so many names The vocabulary looks heavy for a tool whose first example is five lines long. It pays off the moment something goes wrong. When an answer is bad, the cause is in exactly one place: the Document was loaded wrong, the Nodes were split wrong, the retriever fetched the wrong Nodes, or the synthesiser was given the right Nodes and still wrote nonsense. Four suspects, each inspectable on its own. Without the names, you have one black box.
Try it
  1. On paper, write the six nouns in the order data flows through them.
  2. Next to each, write the class name you would use and the method that produces the next noun.
  3. Circle the one place where the number of objects increases.
a sketch where the Document-to-Node step is circled, because that is where twenty files become two thousand chunks. Everything about cost, latency and answer quality is decided there.

Installing it, and the trap in the install docs

LlamaIndex needs Python 3.10 to 3.14. Python 3.9 was dropped during the 0.14 line and will not work. Always install into a virtual environment; the package pulls in a lot of transitive dependencies and you do not want them in your system Python.

BASH
python3 -m venv .venv
source .venv/bin/activate
pip install -U pip
pip install llama-index llama-index-readers-file
export OPENAI_API_KEY="sk-..."

On Windows PowerShell:

POWERSHELL
py -m venv .venv
.venv\Scripts\Activate.ps1
python -m pip install -U pip
pip install llama-index llama-index-readers-file
$env:OPENAI_API_KEY = "sk-..."

In cmd.exe, activate with .venv\Scripts\activate.bat and set the key with set OPENAI_API_KEY=sk-.... To make it survive a new shell, use setx OPENAI_API_KEY "sk-...", which affects future shells only. If you use uv, the equivalent is uv venv && source .venv/bin/activate && uv pip install llama-index llama-index-readers-file.

The second package is not optional, whatever the docs say The official installation page still describes the llama-index bundle as containing llama-index-readers-file. It does not. That dependency was dropped in 0.14.18, and two more (llama-index-indices-managed-llama-cloud, llama-index-readers-llama-parse) in 0.14.19. Without llama-index-readers-file, SimpleDirectoryReader logs `llama-index-readers-file` package not found, some file readers will not be available if not provided by the `file_extractor` parameter. and then reads your PDFs and DOCX files as raw bytes interpreted as text. You get an index full of garbage and no exception. Install both packages.

What the llama-index bundle actually contains as of 0.14.25 is llama-index-core, llama-index-embeddings-openai, llama-index-llms-openai and nltk. That tells you something useful: the bundle is opinionated towards OpenAI. If you are not using OpenAI, do not install the bundle at all.

Verifying the install

BASH
python -c "import llama_index.core as c; print(c.__version__)"
pip show llama-index-core llama-index-workflows
python -c "from llama_index.core.agent.workflow import FunctionAgent; print('ok')"

The first command should print 0.14.25. The third is worth running because it is the import that breaks when you have followed an old tutorial: if it succeeds, your installation has the modern, workflow-based agents.

On Linux and macOS, pip list | grep -i llama shows every LlamaIndex package you have; on Windows use pip list | findstr llama.

Installing without OpenAI

If you have no OpenAI account, or your employer will not let documents leave the building, run both models locally. You need two of them: one LLM to write answers and one embedding model to turn text into vectors.

BASH
pip install llama-index-core llama-index-readers-file llama-index-llms-ollama llama-index-embeddings-huggingface
ollama pull llama3.1
settings_local.py
from llama_index.core import Settings
from llama_index.llms.ollama import Ollama
from llama_index.embeddings.huggingface import HuggingFaceEmbedding

Settings.llm = Ollama(model="llama3.1", request_timeout=360.0, context_window=8000)
Settings.embed_model = HuggingFaceEmbedding(model_name="BAAI/bge-base-en-v1.5")

The docs note that running llama3.1 8B locally wants roughly 32 GB of RAM. The embedding model is far lighter and downloads once to a local cache. Ollama has its own guide in this catalogue if you want the model-serving side in depth.

Note the request_timeout=360.0. Local models are slow, and the default timeouts assume a hosted API. Note also that you must set both Settings.llm and Settings.embed_model. Setting only one is the most frequent beginner mistake with local setups, and the next section explains exactly what it costs you.

Try it
  1. Create a fresh virtual environment and install both packages from the first snippet.
  2. Run all three verification commands and confirm the version is 0.14.25 or later.
  3. Run pip list | grep -i llama and read the list. Identify which package provides the OpenAI LLM.
four or five llama-index-* packages, with llama-index-llms-openai among them. Knowing that provider integrations are separate packages named llama-index-<type>-<name> makes every future ImportError self-solving.

Your first application, in five lines

Make a folder, put something in it, and point LlamaIndex at it.

BASH
mkdir -p rag-demo/data && cd rag-demo
curl -o data/essay.txt https://raw.githubusercontent.com/run-llama/llama_index/main/docs/docs/examples/data/paul_graham/paul_graham_essay.txt

Any text file will do; use your own if you prefer. Now the program:

starter.py
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader

documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents, show_progress=True)
query_engine = index.as_query_engine()
print(query_engine.query("What did the author work on before college?"))

Run it with python starter.py. Expect a progress bar for the embedding step and then a paragraph of prose. This is a complete RAG application, and it is worth being precise about what those five lines did, because each one maps onto a noun from the previous section.

  1. LoadSimpleDirectoryReader("data").load_data() walked the data folder, read each file with a reader chosen by extension, and returned a list of Document objects. By default it does not recurse into subfolders.
  2. Splitfrom_documents applied the default transformations. SentenceSplitter cut each Document into Nodes of about 1024 tokens with 200 tokens of overlap between neighbours.
  3. EmbedEach Node's text went to the embedding model and came back as a vector. This is the step that costs money and takes time, which is why you asked for a progress bar.
  4. StoreThe vectors and the Node text went into the default in-memory store. Nothing is on disk yet.
  5. Queryas_query_engine() built a retriever plus a synthesiser. .query() embedded your question, found the two most similar Nodes, pasted them into a prompt with your question, and sent it to the LLM.

Two defaults in there deserve to be dragged into the light, because they explain most first-day disappointments.

You retrieved two chunks, not ten. similarity_top_k defaults to 2. For a question whose answer lives in one paragraph that is plenty. For "summarise the main themes" it is hopeless, and the model will dutifully summarise the two paragraphs it was given as though they were the whole document. Raise it: index.as_query_engine(similarity_top_k=5).

You used models you never chose. With no Settings configured, the LLM resolves to OpenAI(), whose default model is gpt-3.5-turbo, and the embedding model to OpenAIEmbedding(), whose default is text-embedding-ada-002. Both are old. Always set them explicitly.

Seeing what was actually retrieved

The response object is not just a string. Print its sources and you can audit every answer.

inspect_sources.py
response = query_engine.query("What did the author work on before college?")
print(response.response)
for node in response.source_nodes:
    print(f"--- score={node.score:.3f} file={node.metadata.get('file_name')}")
    print(node.text[:300].replace("\n", " "))

response.source_nodes is a list of NodeWithScore. The score is the similarity between the question's vector and the Node's vector, so higher is better, and the absolute value depends on the embedding model. Get into the habit of printing this. When an answer is wrong, these few lines tell you immediately whether retrieval failed (the right text is not in the list) or synthesis failed (the right text is there and the model still got it wrong). Those two failures have completely different fixes, and guessing which one you have is the most common waste of an afternoon in this field.

Try it
  1. Run starter.py, then ask it the four questions you wrote in the first Try it.
  2. Add the source-printing loop and re-run the question whose answer is not in the documents.
  3. Look at the scores on that question. Did the engine still answer confidently?
two irrelevant chunks with mediocre scores, and often an answer anyway. Retrieval always returns something: there is no built-in notion of "nothing matched". You add that yourself with a similarity cutoff, and knowing it is your job is the point of this exercise.

Configuration: Settings and local overrides

Choosing models explicitly is the first thing to do in any real project. LlamaIndex gives you two levels: a process-wide singleton called Settings, and per-call arguments that override it.

config.py
from llama_index.core import Settings
from llama_index.llms.openai import OpenAI
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.core.node_parser import SentenceSplitter

Settings.llm = OpenAI(model="gpt-4o-mini", temperature=0.1)
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small", embed_batch_size=100)
Settings.text_splitter = SentenceSplitter(chunk_size=512, chunk_overlap=50)
Settings.context_window = 4096
Settings.num_output = 256

Import this module before you build an index and everything downstream picks up these choices. temperature=0.1 is deliberate: for factual question answering you want the model boring and repeatable, not creative.

The attributes Settings holds are llm, embed_model, text_splitter (also reachable as node_parser), chunk_size, chunk_overlap, transformations, tokenizer, callback_manager, context_window and num_output. The verified defaults are worth memorising:

Setting Default in 0.14.25
Settings.llm OpenAI() → gpt-3.5-turbo
Settings.embed_model OpenAIEmbedding() → text-embedding-ada-002
Settings.chunk_size / chunk_overlap 1024 / 200 tokens
similarity_top_k 2
Context window / output when the model is unknown 3900 / 256 tokens
Tokenizer tiktoken cl100k
Persist directory ./storage
Override locally, not globally, when it is a one-off Settings is a process-wide singleton, so changing it affects everything. Prefer local arguments when the change belongs to one call: VectorStoreIndex.from_documents(docs, embed_model=...), index.as_query_engine(llm=..., similarity_top_k=5), SentenceSplitter(chunk_size=256) passed in transformations=[...]. A habit that saves real pain later: in a web server, mutating Settings per request is a race condition waiting to happen.

One piece of history you will trip over in older tutorials: configuration used to live in a ServiceContext object. It still imports, but using it raises ValueError: ServiceContext is deprecated. Use llama_index.settings.Settings instead, or pass in modules to local functions/methods/interfaces. If a blog post hands you ServiceContext.from_defaults(...), the post predates mid-2024 and the rest of its code is probably stale too.

Chunk size is a real decision

chunk_size controls how much text goes into each Node, and it is the knob with the largest effect on answer quality. Small chunks (256–512 tokens) are precise: a match means the chunk is genuinely about your question, and you can afford to retrieve several. Large chunks (1024–2048) carry more surrounding context, which helps when the answer needs a paragraph of setup, but they dilute the embedding — a 2000-token chunk covering five topics has a vector that is the average of five topics and matches none of them strongly.

chunk_overlap exists so that a sentence straddling a boundary still appears whole in one of the two chunks. Keep it well below the chunk size; the framework enforces this and will tell you off: ValueError: Got a larger chunk overlap (200) than chunk size (100), should be smaller. If you shrink chunk_size, shrink the overlap too.

There is no universally right answer. Start at 512 with 50 overlap for prose documents, measure, and adjust. The measuring part is not optional, and tools like RAGAS exist to make it systematic once you are past the prototype.

Try it
  1. Build the same index three times with chunk_size set to 256, 512 and 2048.
  2. For each, print len(index.docstore.docs) to see how many Nodes you got.
  3. Ask the same narrow factual question of all three with similarity_top_k=3 and compare the retrieved text.
many small precise chunks versus few broad ones, and a visible difference in how much irrelevant text reaches the model. This single experiment teaches more about RAG than any amount of reading.

Loading files properly

SimpleDirectoryReader does more than its name suggests, and its defaults will bite you at least once. The two that matter most: it reads only the top level of the directory (recursive=False), and it assumes UTF-8.

loading.py
from llama_index.core import SimpleDirectoryReader

reader = SimpleDirectoryReader(
    input_dir="data",
    recursive=True,
    required_exts=[".pdf", ".md", ".txt"],
    exclude=["drafts/*"],
    num_files_limit=200,
)
documents = reader.load_data()
print(f"{len(documents)} documents")
print(documents[0].metadata)

The options you will reach for are input_dir, input_files=[...] for an explicit list, exclude=[...], recursive=True, required_exts=[...], num_files_limit, encoding="latin-1" for legacy files, file_metadata=fn to attach your own metadata, file_extractor={".myfile": MyReader()} to override the reader for an extension, and fs= to read from a remote filesystem such as S3. load_data(num_workers=4) parallelises across processes, and iter_data() streams instead of loading everything into memory.

Supported extensions out of the box include .csv, .docx, .epub, .hwp, .ipynb, .jpeg, .jpg, .mbox, .md, .mp3, .mp4, .pdf, .png, .ppt, .pptm and .pptx — all of which need llama-index-readers-file installed. JSON is handled by a separate package, llama-index-readers-json.

Every Document arrives with metadata filled in: file_path, file_name, file_type, file_size, creation_date, last_modified_date and last_accessed_date. The dates are UTC in %Y-%m-%d form. Print it on your first run of any new dataset. If file_name is right but the text looks like %PDF-1.4 followed by binary noise, you have the missing-readers problem from the install section.

Metadata is not free Metadata is prepended to the Node's text before embedding and before the LLM sees it, so it consumes your chunk budget. Attach too much and you get ValueError: Metadata length (N) is longer than chunk size (M). Consider increasing the chunk size or decreasing the size of your metadata to avoid this. You can also exclude keys selectively with excluded_embed_metadata_keys and excluded_llm_metadata_keys on a Document, which is the right fix when you want metadata for filtering but not for matching.

Windows users have two extra notes from the docs. load_data(num_workers=N) uses multiprocessing, so "Windows users may see less or no performance gains", and any multiprocessing code must sit under if __name__ == "__main__": or it will spawn copies of itself. Several Windows encoding bugs were fixed during the 0.14 line by forcing UTF-8 for text-mode file I/O.

Try it
  1. Put a PDF in a subfolder of data and load with the defaults. Count the documents.
  2. Add recursive=True and count again.
  3. Print documents[0].text[:200] and confirm it is readable prose, not binary.
zero or fewer documents on the first run than you expected, because the subfolder was skipped silently. Silent is the operative word: no warning, no error, just a smaller index than you thought you had.

Persisting the index so you build it once

Everything so far lives in RAM and dies with the process. Embedding is the expensive step — in money with a hosted model, in time with a local one — so rebuilding an index on every run is the first thing to fix.

persist.py
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader

documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)
index.storage_context.persist("storage")
reload.py
from llama_index.core import StorageContext, load_index_from_storage

storage_context = StorageContext.from_defaults(persist_dir="storage")
index = load_index_from_storage(storage_context)
print(index.as_query_engine().query("What is this about?"))

persist() with no argument uses ./storage. Look inside the folder afterwards: docstore.json holds the Nodes, index_store.json the index metadata, and default__vector_store.json the embeddings. They are plain JSON, so you can open them and see exactly what your chunking produced. That is a genuinely useful debugging move the first time chunking surprises you.

A StorageContext bundles four or five stores: a docstore for Nodes, an index store for index metadata, a vector store for embeddings, and optionally a property graph store and chat stores. At this level the defaults are all in-memory-plus-JSON, which is right for prototypes and wrong for production.

If you keep several indexes in one directory, name them, or loading becomes ambiguous:

multi_index.py
index.set_index_id("vector_index")
index.storage_context.persist("storage")

# later
index = load_index_from_storage(storage_context, index_id="vector_index")

The three errors you will meet here are all informative. A wrong or missing directory gives FileNotFoundError: [Errno 2] No such file or directory: '.../docstore.json'. An empty store gives ValueError: No index in storage context, check if you specified the right persist_dir. Several unnamed indexes give ValueError: Expected to load a single index, but got N instead. Please specify index_id.

Load if present, build if not The pattern worth writing once and reusing everywhere: check whether the persist directory exists, load if it does, build and persist if it does not. Three lines, and it turns a twenty-second startup into a half-second one for the rest of the project's life.
build_or_load.py
import os
from llama_index.core import (VectorStoreIndex, SimpleDirectoryReader,
                              StorageContext, load_index_from_storage)

PERSIST_DIR = "storage"

if os.path.exists(PERSIST_DIR):
    index = load_index_from_storage(StorageContext.from_defaults(persist_dir=PERSIST_DIR))
else:
    documents = SimpleDirectoryReader("data").load_data()
    index = VectorStoreIndex.from_documents(documents)
    index.storage_context.persist(PERSIST_DIR)

One thing to understand about persistence before you move on: the JSON stores are versioned with the library, so test that your saved index still loads after you upgrade. And re-index from scratch whenever you change the embedding model or the chunking, because old vectors from a different model are meaningless against new queries. When you graduate to a real vector database — Chroma, Qdrant, pgvector — most of them store text alongside vectors and persist() becomes unnecessary; FAISS is the notable exception, because it stores vectors only and still needs a docstore beside it.

Try it
  1. Run build_or_load.py twice and time both runs.
  2. Open storage/docstore.json and find the text of one Node. Check the overlap against its neighbour.
  3. Delete the storage folder and run reload.py to see the FileNotFoundError for yourself.
a dramatic second-run speedup, visible overlap between adjacent Nodes, and one deliberate error you will now recognise instantly in someone else's traceback.

Asking questions well: query engines and chat engines

A query engine answers one question with no memory of the last one. That is the right shape for a search box and the wrong shape for a conversation. Both are one method call away.

querying.py
from llama_index.core.postprocessor import SimilarityPostprocessor

query_engine = index.as_query_engine(
    similarity_top_k=5,
    response_mode="compact",
    node_postprocessors=[SimilarityPostprocessor(similarity_cutoff=0.7)],
)
response = query_engine.query("What are the main themes?")
print(response)

Three choices there, each solving a specific problem. similarity_top_k=5 gets you past the default of two. The SimilarityPostprocessor drops Nodes whose score is below 0.7, which is how you stop the engine answering from irrelevant text — the gap you found in the fourth Try it. And response_mode picks the synthesis strategy.

The synthesis modes you should know are three. compact, the default for query engines, stuffs as many retrieved Nodes as fit into one prompt and asks once, then repeats with the leftovers if there are any. It is the cheapest and right for most questions. refine sends the first Node with the question, then shows the model its own draft answer plus the next Node and asks it to improve the answer, one Node at a time. More calls, more cost, better on questions whose answer is assembled from several places. tree_summarize summarises Nodes in pairs up a tree until one summary remains, which is the mode for "summarise everything".

Pick the mode from the question's shape. A lookup question wants compact. A whole-corpus question wants tree_summarize and a much higher top_k. Getting this wrong is a common cause of "the summary only mentioned two sections".

Streaming

Waiting eight seconds in silence feels broken even when it is not. Stream instead:

streaming.py
query_engine = index.as_query_engine(streaming=True, similarity_top_k=5)
streaming_response = query_engine.query("Summarize the documents.")
streaming_response.print_response_stream()

Multi-turn conversation

chatting.py
chat_engine = index.as_chat_engine()
print(chat_engine.chat("What does the document say about pricing?"))
print(chat_engine.chat("And what about discounts?"))

for token in chat_engine.stream_chat("Summarise that in one sentence.").response_gen:
    print(token, end="", flush=True)

The second question has no subject — "what about discounts" is meaningless alone — and it works because the chat engine carries history. Since 0.13 the default chat engine is CondensePlusContextChatEngine. The name describes the mechanism: it condenses your new message plus the history into a single standalone question, retrieves with that, then puts the retrieved context into the prompt along with the history. Understanding this explains the main failure mode: if the condensed question drops a crucial detail from three turns ago, retrieval goes looking for the wrong thing, and the answer is confidently about the wrong topic.

A stale docstring you will meet Before 0.13, as_chat_engine() returned an agent-based engine by default. That class was removed. The default chat_mode=ChatMode.BEST now maps to condense-plus-context — but the docstring in the source still claims BEST "uses an agent". The docstring is wrong and the behaviour is what matters. If you need agentic behaviour over your index, build an agent explicitly, as the next section does.
Try it
  1. Ask the same broad question with response_mode="compact" and then "tree_summarize", both with similarity_top_k=10.
  2. Add SimilarityPostprocessor(similarity_cutoff=0.8) and re-ask your out-of-scope question.
  3. Have a four-turn conversation with the chat engine where the fourth turn depends on the first.
a broader summary from tree_summarize, a refusal or a much shorter answer once the cutoff is on, and at least one conversational turn where the engine loses the thread. That last one is not a bug you can fix with a flag; it is the condense step's limit.

Your first agent

A query engine looks things up. An agent decides what to do. The difference matters when a question needs arithmetic, an API call, or a choice between two sources.

Agents in LlamaIndex are asynchronous. There is no synchronous agent.chat() — it was removed in 0.13 along with the whole previous generation of agent classes. Every run is await agent.run(...), which means your code needs an async entry point.

first_agent.py
import asyncio
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI


def multiply(a: float, b: float) -> float:
    """Useful for multiplying two numbers."""
    return a * b


def add(a: float, b: float) -> float:
    """Useful for adding two numbers together."""
    return a + b


agent = FunctionAgent(
    tools=[multiply, add],
    llm=OpenAI(model="gpt-4o-mini"),
    system_prompt="You are a helpful assistant that can do arithmetic.",
)


async def main():
    print(await agent.run("What is 1234 * 4567, then add 100?"))


if __name__ == "__main__":
    asyncio.run(main())

Notice what you did not write: no schema, no JSON, no registration. You passed plain Python functions and LlamaIndex derived the tool definitions. The rules are worth stating explicitly because they are the whole interface:

  • The function name becomes the tool name.
  • The docstring becomes the description, which is what the model reads when deciding whether to call it. A vague docstring is the single most common cause of an agent ignoring a perfectly good tool.
  • Type hints become the parameter schema. Untyped parameters give the model nothing to go on. You can add per-parameter descriptions with Annotated[str, "the city name"].

For more control, wrap the function yourself: FunctionTool.from_defaults(fn, async_fn=..., name=..., description=..., return_direct=...). Setting return_direct=True ends the agent loop immediately and returns the tool's output as the final answer, which is useful when the tool already produces exactly what the user asked for and you do not want to pay for the model to paraphrase it.

Giving the agent your documents

The interesting move is combining the two halves of this guide: wrap a query engine as a tool, and the agent can search your documents when — and only when — it decides it needs to.

rag_agent.py
import asyncio
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.core.tools import QueryEngineTool
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI

documents = SimpleDirectoryReader("data").load_data()
index = VectorStoreIndex.from_documents(documents)

docs_tool = QueryEngineTool.from_defaults(
    query_engine=index.as_query_engine(similarity_top_k=5),
    name="company_handbook",
    description=(
        "Answers questions about the company handbook: leave policy, "
        "working hours, expenses and travel. Input is a natural-language question."
    ),
)

agent = FunctionAgent(
    tools=[docs_tool],
    llm=OpenAI(model="gpt-4o-mini"),
    system_prompt="Answer from the handbook tool. If it has no answer, say so plainly.",
)


async def main():
    print(await agent.run("How many days of annual leave do I get?"))


if __name__ == "__main__":
    asyncio.run(main())

Write that description as if it were documentation for a colleague who can only see the description, because that is exactly the situation the model is in. "Searches documents" tells it nothing. Naming the topics and saying what the input looks like is what makes routing work, and with several tools it is the difference between an agent that picks the right source and one that guesses.

Memory across runs

Each agent.run() is independent by default. To carry a conversation, create a Context and pass it every time.

agent_memory.py
import asyncio
from llama_index.core.workflow import Context
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.llms.openai import OpenAI

agent = FunctionAgent(tools=[], llm=OpenAI(model="gpt-4o-mini"))


async def main():
    ctx = Context(agent)
    print(await agent.run("My name is Logan.", ctx=ctx))
    print(await agent.run("What is my name?", ctx=ctx))


if __name__ == "__main__":
    asyncio.run(main())

For anything longer-lived, Memory.from_defaults(session_id="user-42") gives you a proper memory object with a token budget (30,000 tokens by default, in-memory SQLite storage) that you pass as memory= to run(). The older ChatMemoryBuffer class is documented as deprecated in favour of Memory, though agents still construct one internally when you pass no memory at all. Use Memory in new code.

Two agent limits will eventually interrupt you. The iteration cap is 20: exceed it and you get WorkflowRuntimeError: Max iterations of 20 reached! Either something went wrong, or you can increase the max iterations with .run(.., max_iterations=...) or use early_stopping_method='generate' to generate a final response instead. When you see that, suspect a loop — a tool that keeps failing and a model that keeps retrying it — before you raise the cap. Second, if your model does not support streaming, pass streaming=False, since agents stream by default.

ReActAgent, imported from the same module, is the alternative when your model has no native tool-calling API: it prompts for Thought/Action/Observation text and parses it. CodeActAgent writes and executes code and requires you to supply a code_execute_fn — it raises ValueError: code_execute_fn must be provided for CodeActAgent otherwise. Do not hand it a bare exec; running model-written code needs a sandbox, which is a Senior-level topic for a good reason. If multi-agent orchestration is where you are heading, LangGraph covers the graph-shaped approach to the same problem.

Try it
  1. Run first_agent.py and then deliberately replace multiply's docstring with "Does a thing.". Re-run.
  2. Build rag_agent.py over your own documents and ask it something out of scope.
  3. Add a second tool with an overlapping description and see which one the agent picks.
an agent that stops using the tool once its docstring is useless, and visible confusion when two descriptions overlap. Tool descriptions are prompt engineering, and this is the cheapest way to learn it.

Reading the errors

LlamaIndex has unusually informative error messages. Learning to read six of them will save you more time than any other half-hour in this guide.

No API key. A long ValueError beginning Could not load OpenAI model. If you intended to use OpenAI, please check your OPENAI_API_KEY. The embedding variant says Could not load OpenAI embedding model. ... Consider using embed_model='local'. The cause is that something resolved to OpenAI because Settings left it unset. The sneaky case is a local setup where you set Settings.llm = Ollama(...) and forgot Settings.embed_model: the LLM is local, the embedder silently is not, and the error appears at index time rather than query time. Set both.

Missing integration package. ImportError: `llama-index-llms-openai` package not found, please run `pip install llama-index-llms-openai`, or a plainer ModuleNotFoundError: No module named 'llama_index.llms.ollama'. The naming rule makes this self-solving: a package called llama-index-{type}-{name} imports as llama_index.{type}.{name}. Read the failed import backwards and you have the pip command.

Legacy imports from old tutorials. ImportError: cannot import name 'FunctionCallingAgent' from 'llama_index.core.agent' or ModuleNotFoundError: No module named 'llama_index.core.query_pipeline', or an AttributeError on ReActAgent.from_tools or agent.chat. All of these mean the tutorial predates 0.13, which removed the old agent classes and QueryPipeline entirely. Modern code imports FunctionAgent and ReActAgent from llama_index.core.agent.workflow and uses await agent.run(...). There is no compatibility shim, so a tutorial that hits this needs replacing, not patching.

Embedding dimension mismatch. ValueError: shapes (384,) and (1536,) not aligned: 384 (dim 0) != 1536 (dim 0). You built the index with one embedding model and queried it with another — commonly a 384-dimension local model against an index built with a 1536-dimension OpenAI one. Vectors from different models are not comparable. Rebuild the index, or query with the model you built it with. Hosted vector databases raise their own version of this, usually at collection-creation time.

Prompt too big. ValueError: Calculated available context size -123 was not non-negative. The arithmetic is context window minus prompt minus reserved output, and it went negative. Either the model's context window is unknown — so the framework used its 3900-token fallback — or your local model genuinely has a small one. Fix it by setting a correct context_window and num_output on the LLM or in Settings, or by lowering similarity_top_k or chunk_size. A 2048-token chunk times top_k=10 is 20,000 tokens of context; the arithmetic is not mysterious once you do it.

Async misuse. RuntimeError: asyncio.run() cannot be called from a running event loop appears when you call asyncio.run() inside a Jupyter notebook or a FastAPI handler, both of which already run a loop. In a notebook, just await agent.run(...) directly. In a script, have exactly one asyncio.run(main()). The related SyntaxError: 'await' outside function means you pasted agent code into a script without wrapping it in async def main().

Do not use the CLI You may find tutorials using llamaindex-cli rag, llamaindex-cli upgrade, download_llama_pack or download_loader. The llama-index-cli package is deprecated and unmaintained, it was dropped from the umbrella bundle in 0.14.20, and on current core it prints DeprecationWarning: llama-index-cli is deprecated and no longer maintained and then crashes with ModuleNotFoundError: No module named 'llama_index.core.download'. There is no supported CLI for the framework. Everything is Python.
Try it
  1. Unset OPENAI_API_KEY in a scratch shell and run starter.py. Read the whole error.
  2. Build an index with HuggingFaceEmbedding, persist it, then load and query it with the OpenAI default.
  3. Paste await agent.run("hi") into a bare script and run it.
three errors you caused on purpose. Each one costs an hour the first time it surprises you and thirty seconds once you have met it deliberately.

Putting it all together

One small project that uses everything above: a command-line assistant over a folder of documents that builds its index once, caches it, retrieves with a similarity floor, and runs as an agent so it can refuse out-of-scope questions instead of inventing answers.

assistant.py
import asyncio
import os
import sys

from llama_index.core import (Settings, SimpleDirectoryReader, StorageContext,
                              VectorStoreIndex, load_index_from_storage)
from llama_index.core.agent.workflow import FunctionAgent
from llama_index.core.node_parser import SentenceSplitter
from llama_index.core.postprocessor import SimilarityPostprocessor
from llama_index.core.tools import QueryEngineTool
from llama_index.core.workflow import Context
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.llms.openai import OpenAI

PERSIST_DIR = "storage"
DATA_DIR = "data"

Settings.llm = OpenAI(model="gpt-4o-mini", temperature=0.1)
Settings.embed_model = OpenAIEmbedding(model="text-embedding-3-small")
Settings.text_splitter = SentenceSplitter(chunk_size=512, chunk_overlap=50)


def build_or_load_index() -> VectorStoreIndex:
    if os.path.exists(PERSIST_DIR):
        print(f"loading index from {PERSIST_DIR}")
        return load_index_from_storage(
            StorageContext.from_defaults(persist_dir=PERSIST_DIR)
        )
    print(f"building index from {DATA_DIR}")
    documents = SimpleDirectoryReader(
        DATA_DIR, recursive=True, required_exts=[".pdf", ".md", ".txt"]
    ).load_data()
    if not documents:
        sys.exit(f"no documents found in {DATA_DIR}")
    print(f"loaded {len(documents)} documents")
    index = VectorStoreIndex.from_documents(documents, show_progress=True)
    index.storage_context.persist(PERSIST_DIR)
    return index


def build_agent(index: VectorStoreIndex) -> FunctionAgent:
    query_engine = index.as_query_engine(
        similarity_top_k=5,
        response_mode="compact",
        node_postprocessors=[SimilarityPostprocessor(similarity_cutoff=0.7)],
    )
    tool = QueryEngineTool.from_defaults(
        query_engine=query_engine,
        name="document_search",
        description=(
            "Searches the local document collection and answers questions "
            "about its contents. Input is a full natural-language question."
        ),
    )
    return FunctionAgent(
        tools=[tool],
        system_prompt=(
            "You answer questions using the document_search tool only. "
            "If the tool returns nothing relevant, say you could not find it "
            "in the documents. Never guess."
        ),
    )


async def main() -> None:
    agent = build_agent(build_or_load_index())
    ctx = Context(agent)
    print("ask a question, or an empty line to quit")
    while True:
        try:
            question = input("\n> ").strip()
        except (EOFError, KeyboardInterrupt):
            break
        if not question:
            break
        print(await agent.run(question, ctx=ctx))


if __name__ == "__main__":
    asyncio.run(main())

Read it once against the six nouns. SimpleDirectoryReader produces Documents. SentenceSplitter, configured on Settings, turns them into Nodes. VectorStoreIndex.from_documents embeds and indexes them. persist and load_index_from_storage make that work survive a restart. as_query_engine assembles a retriever, a postprocessor and a synthesiser. QueryEngineTool presents the whole thing to the agent as one capability, and Context keeps the conversation coherent across turns.

Four exercises turn this from a demo into something you understand. First, ask a question you know is out of scope and verify that the system prompt plus the similarity cutoff actually produce a refusal — if it still invents an answer, raise the cutoff and look at the scores to decide where the real threshold is. Second, add a second tool, a plain Python function, and watch the agent choose. Third, delete storage, change chunk_size to 1024, rebuild, and compare answers on the same five questions. Fourth, swap both models for local ones using the Ollama snippet from the install section and confirm that nothing else in the file has to change — that substitutability is the payoff for learning the vocabulary.

Habits that hold up

  • Both Settings.llm and Settings.embed_model set explicitly
  • similarity_top_k chosen for the question, not left at 2
  • A similarity cutoff, so "no match" is possible
  • Index persisted and loaded, never rebuilt per run
  • response.source_nodes printed while debugging
  • Tool descriptions written for a reader who sees nothing else

What goes wrong

  • Defaults left in place, so you ship gpt-3.5-turbo and ada-002
  • Two chunks retrieved for a whole-corpus question
  • Confident answers from irrelevant text
  • Re-embedding the corpus on every run
  • Guessing whether retrieval or synthesis failed
  • "Searches documents" as a tool description

What you can now do, and what comes next

You can install LlamaIndex without falling into the readers-file trap, load a folder of mixed file types, control how it is chunked, build and persist a vector index, query it with a sensible top_k and a similarity floor, hold a multi-turn conversation over it, inspect exactly which chunks produced an answer, wrap the whole thing as a tool for an agent, give that agent your own Python functions, and read the framework's six most common errors without reaching for a search engine. You also know which tutorials to throw away: anything with ServiceContext, QueryPipeline, FunctionCallingAgent, agent.chat() or llamaindex-cli is from before 0.13 and will not run.

What you have deliberately not done is anything that survives contact with scale. The in-memory vector store loads entirely into RAM and runs in one process. Re-ingesting a changed folder re-embeds everything, including the 99% that did not change. Nothing measures whether answers are actually correct. There is no tracing, so when quality drops in production you have no record of what was retrieved. And the chunking is uniform, which is the wrong strategy for documents full of tables.

Mid-level picks up precisely there: IngestionPipeline with a docstore and a cache so unchanged documents are skipped, real vector stores and metadata filters, retrieval tuning and reranking, memory blocks, multi-agent workflows with AgentWorkflow, evaluation, and the async patterns that let one process serve many users. Senior goes on to custom Workflows, durable execution and human-in-the-loop, multi-tenancy, cost control and the security model — which matters because LlamaIndex is a library with no built-in authentication, authorisation, tenancy or sandboxing, and every one of those is yours to supply.

Three directions from here, depending on where you are going. For retrieval quality, pair this with a real vector database — Qdrant, Chroma or pgvector — and learn metadata filtering properly. For trustworthiness, add tracing with Langfuse and scoring with RAGAS; an untraced, unevaluated RAG system is one you cannot improve, because you cannot tell whether a change helped. For serving, put the whole thing behind FastAPI, which is the natural fit for an async-first framework. And if you want to compare approaches before committing, read the LangChain and LangGraph guides; knowing why you chose one is worth more in an interview than knowing one deeply.

One last habit for the long run. Pin your versions — llama-index-core>=0.14.25,<0.15 — because the minor number is the breaking-change boundary in this project, and every integration package declares that same <0.15 bound. When 0.15 lands, all of them will be re-released together, and the upgrade will be a deliberate afternoon rather than a surprise on a Monday morning. Read the changelog before you take it.

Sources