This is part one of three. By the end of it you can install Haystack, build a pipeline out of typed components, index a folder of your own documents, ask questions about them with a language model, save the whole thing to a YAML file and load it back, and read the error messages Haystack produces well enough to fix them yourself. Nothing here is a teaser for the paid version of anything: Haystack is an open-source Python library, and everything below runs on a laptop.
Each section ends with a Try it task. Do them as you go. Pipelines are one of those ideas that look obvious on a page and then behave unexpectedly the first time you wire two components together, and the only cure is to have wired two components together.
One thing to settle before the first line of code. This guide targets Haystack 3.2.0, the current stable release. Haystack 3.0 arrived in July 2026 and removed a long list of things that 2.x code used daily. Most blog posts, most Stack Overflow answers and most of what a language model will confidently tell you about Haystack describe 2.x. If you copy an example that imports AsyncPipeline or OpenAIGenerator, it will not run, and the error will not explain why. Section after section below flags the places where 2.x habits break, because recognising them is most of what separates a frustrating first week from a productive one.
What Haystack is, and the problem it solves
Haystack is a Python framework from deepset for building applications on top of language models: question answering over your own documents, search, chatbots, agents that call tools. Its central claim is small and concrete. Instead of writing one long script that fetches documents, builds a prompt string, calls an API and parses the reply, you build a graph of small typed parts, and the framework checks that the parts fit together before you run anything.
To see why that matters, write the long script first, in your head. A retrieval-augmented question answering app — the standard first project in this field, usually shortened to RAG, for retrieval-augmented generation — does roughly this: take the user's question, find the handful of passages in your corpus most likely to contain the answer, paste those passages into a prompt along with the question, send that prompt to a language model, and return the model's reply. Five steps, maybe sixty lines.
The sixty lines are fine until you need to change something. Swap the keyword search for a vector search and the retrieval step's inputs change shape. Add a reranking step between retrieval and the prompt and you are reordering function calls and hoping you passed the right variable. Serve it behind a web API and you discover the model client was built at import time with a key that was not set yet. Run the same app against a second corpus and you copy the file. Each of these is a small problem, and they compound into the thing that makes LLM codebases unpleasant: nothing has a declared shape, so nothing can be checked, so every change is a guess verified by running it.
Haystack's answer is to make the shape explicit. A component declares what types it accepts and what types it emits. A pipeline is a set of components plus a set of connections between their named inputs and outputs, and connections are type-checked when you create them, not when you run the pipeline. If you connect a component that emits list[Document] to one that expects str, you get an error at wiring time, with both type names in the message, before you have spent a single token of API budget.
The second half of the claim is that the graph is data. A pipeline can be serialised to YAML and loaded back, which means the structure of your application can live in a file that is reviewed, versioned and deployed separately from the code that runs it. That is an unusual property for this kind of library and it is the reason Haystack shows up in teams that already have opinions about configuration management.
What came before is worth knowing, because it explains some of Haystack's shape. Haystack 1.x, released long before the current wave of chat models, was built around extractive question answering: a reader model pointed at a span of text inside a retrieved document and said "the answer is these words, here". It had Node classes, a YAML-first design and a quite rigid pipeline. Haystack 2.x, a full rewrite, replaced that with the component-and-socket model you will learn here and made generative models first-class. Haystack 3.x keeps the 2.x model and cleans up the parts that had accumulated duplication — most visibly by deleting the old non-chat generators and the separate asynchronous pipeline class. The 1.x package is still on PyPI under the name farm-haystack, which matters for one practical reason covered in the install section: it claims the same haystack import name, so the two cannot live in one environment.
FutureWarning and is gone in 3.2, which gives you about a month. Pin your dependency, read the upgrade notes, and never ignore a FutureWarning from this library.
- Write down the five steps of the RAG flow described above, in order, on paper.
- Next to each step, write the Python type you think flows out of it.
- Keep the paper. By the end of this guide you will have built exactly that graph, and you can check how close your guesses were.
Document, and it carries the metadata that makes citations possible.
The four nouns
Almost everything in Haystack is one of four things. Learn these names properly and the API documentation stops being a maze.
A Component is a Python class decorated with @component that has a run() method returning a dictionary. The parameters of run() are its inputs. The keys of the returned dictionary are its outputs, declared ahead of time with the @component.output_types(...) decorator. Each named input and output is called a socket, and a socket has a type. That is the whole contract. A component can be used entirely on its own — call run() yourself — or dropped into a pipeline.
A Pipeline is a directed multigraph of components. "Multigraph" matters: two components can be joined by more than one connection, because connections are between sockets, not between components. Data travels only along the connections you declare. A component sitting in the pipeline with nothing wired into it does not magically receive the inputs you passed to its neighbour. This is the single most common early misunderstanding, and the cure is to read the graph as a set of wires rather than a sequence of steps.
A Document is a dataclass holding a piece of content plus everything you know about it: id, content, blob for binary data, meta for your own metadata dictionary, score for relevance after retrieval, and embedding and sparse_embedding for vectors. Documents are the currency of the retrieval half of Haystack. Converters produce them, splitters cut them up, embedders add vectors to them, retrievers return them with scores attached, and prompt builders loop over them to build text.
A Document Store is a database interface — and it is not a component. It has no run() method and cannot be added to a pipeline. Its protocol has exactly four methods: count_documents, filter_documents, write_documents and delete_documents. Components that need storage hold a reference to one. So a retriever takes document_store=... in its constructor, and the thing you add to the pipeline is the retriever. Beginners try to add the store and get a confusing error; now you know why.
Two more words you will meet within the hour. A ChatMessage is a message with a role (system, user, assistant or tool) and content parts; you build one with ChatMessage.from_user(...), ChatMessage.from_system(...) and friends, and you read its text with .text. A Chat Generator is the component that actually calls a language model, for example OpenAIChatGenerator. It takes messages and returns {"replies": [ChatMessage, ...]}.
That diagram is the shape of nearly every Haystack application you will build this year. The top lane runs occasionally and writes into the store. The bottom lane runs on every request and reads from it. They are separate pipelines that happen to share one store object, and keeping them separate is a deliberate choice rather than an accident: indexing is slow, batchy and idempotent, while querying is fast and latency-sensitive.
There is also a lifecycle worth naming now because it explains an error you will hit. Components acquire expensive resources — model weights, HTTP clients — in warm_up(), not in __init__. Haystack calls warm_up() automatically before the first run, and close() releases the resources. The consequence: constructing a chat generator with no API key set succeeds silently, and the failure appears later, at warm-up or on the first run. In Haystack 2.x it failed at construction. If you remember that change you will save yourself a confused ten minutes.
- Say out loud which of these four is not a component: Document Store, Retriever, DocumentWriter, DocumentSplitter.
- For the retriever, name the two things it needs: one at construction time, one at run time.
pipeline.run() call.
Installing Haystack and checking it works
Haystack needs Python 3.10 or later. The package on PyPI is called haystack-ai, which is a common tripwire: the import name is haystack but the install name is not.
Use a virtual environment. Not as a matter of taste — Haystack's optional dependencies are numerous and you will install several of them experimentally, and you want that mess scoped to one folder.
python3 -m venv .venv
source .venv/bin/activate
pip install haystack-ai
On Windows PowerShell the equivalent is three commands:
py -m venv .venv
.venv\Scripts\Activate.ps1
pip install haystack-ai
If your team standardises on uv, either form works:
uv pip install haystack-ai # as an installer, inside an existing environment
uv add haystack-ai # as a project dependency, writing to pyproject.toml
And for conda users:
conda install conda-forge::haystack-ai
Now verify. Do not trust pip install output; ask the library what it thinks its own version is.
python -c "from haystack.version import __version__; print(__version__)"
3.2.0
If that prints something starting with 2., your environment has an older release and roughly a third of the code in this guide will fail. Upgrade with pip install --upgrade haystack-ai — the package name did not change between 2.x and 3.x, only the contents. pip show haystack-ai gives you the full picture including where it was installed, which is useful when you suspect you are in the wrong environment.
farm-haystack is Haystack 1.x. It uses the same haystack import namespace, so installing both leaves you with a half-overwritten package and import errors that make no sense. If you inherit a 1.x project, migrate it in steps — 1.x to 2.x, then 2.x to 3.x — in separate environments. The 1.x documentation is archived as a downloadable ZIP, linked from the official FAQ.
There is a Docker image if you prefer not to install Python things on your laptop at all. The only flavour is base, which is equivalent to pip install haystack-ai and nothing more.
docker pull deepset/haystack:base-v3.2.0
docker run -it --rm deepset/haystack:base-v3.2.0 \
python -c "from haystack.version import __version__; print(__version__)"
To add integrations to that image, extend it:
FROM deepset/haystack:base-v3.2.0
RUN pip install sentence-transformers-haystack qdrant-haystack
If containers are new to you, the Docker guide in this series covers that FROM line and everything around it.
Optional dependencies are lazy, and this is a feature. Haystack does not install a PDF parser, an HTML extractor or a tokeniser unless you ask. The first time you use a component that needs one, you get a precise message:
Haystack failed to import the optional dependency 'pypdf'. Run 'pip install pypdf'.
Do what it says. The common ones are pypdf for PDF conversion, trafilatura for HTML extraction, tiktoken for token-based splitting, and rich for the fancy console confirmation UI.
Last piece of setup: an API key, if you want to call a hosted model. OpenAIChatGenerator reads OPENAI_API_KEY from the environment.
export OPENAI_API_KEY="sk-..."
$env:OPENAI_API_KEY="sk-..." # this session only
setx OPENAI_API_KEY "sk-..." # persisted for future sessions
Two related environment variables are worth knowing early because they explain hangs: OPENAI_TIMEOUT defaults to 30 seconds and OPENAI_MAX_RETRIES defaults to 5. A badly behaving endpoint can therefore keep your process busy for well over two minutes before raising.
You can do real work with no key at all. Haystack 3.0 added MockChatGenerator, MockTextEmbedder and MockDocumentEmbedder precisely so that tests and tutorials can exercise pipeline structure without network calls or spend. A keyword-search pipeline over InMemoryDocumentStore also needs nothing external. Both appear below.
- Create a fresh virtual environment and install
haystack-ai. - Print the version and confirm it is 3.2.0 or later.
- Run
python -c "from haystack.components.generators import OpenAIGenerator"and read the error.
OpenAIGenerator was removed in 3.0 along with the other non-chat generators. Seeing this failure deliberately, once, saves you from mistaking it for a broken install later.
Your first component
Before pipelines, build a component on its own. It is ten lines and it makes the contract concrete.
from haystack import component
@component
class WelcomeTextGenerator:
@component.output_types(welcome_text=str, note=str)
def run(self, name: str):
return {"welcome_text": f"Hello {name}".upper(), "note": "ready"}
greeter = WelcomeTextGenerator()
print(greeter.run(name="Layla"))
{'welcome_text': 'HELLO LAYLA', 'note': 'ready'}
Read every line of that against the rules from the previous section. The @component decorator registers the class. run() takes one parameter, name, typed as str — that is one input socket called name. The decorator declares two output sockets, welcome_text and note, both strings. The returned dictionary's keys match the declared outputs exactly; if they do not, Haystack raises rather than silently dropping data. And because the class is a normal Python object, you can call run() directly, which makes components genuinely unit-testable without a pipeline anywhere in sight.
Three rules about components that will save you debugging time later. Do not mutate your inputs. If you receive a list of documents and want to change them, copy the list or use dataclasses.replace() on each document; the same objects may be handed to another component. Store every constructor argument on an attribute of the same name — self.model = model for a model= parameter — because the default serialisation walks the constructor signature and reads the matching attributes, and a mismatch produces an error at save time rather than a wrong file. And if you add an asynchronous version, it must be called run_async, must be a coroutine, and must declare exactly the same parameters and output types as run; Haystack checks all three and the error messages say so plainly.
- Copy
welcome.pyand run it. - Change the returned dictionary key from
"note"to"notes"and run again. - Remove the
@component.output_typesdecorator entirely and run again.
Documents, stores and a first pipeline
Now the retrieval half. Start with the simplest store there is.
from haystack import Document
from haystack.document_stores.in_memory import InMemoryDocumentStore
store = InMemoryDocumentStore()
store.write_documents([
Document(content="Riyadh is the capital of Saudi Arabia.", meta={"topic": "geography"}),
Document(content="Cairo sits on the Nile and is the capital of Egypt.", meta={"topic": "geography"}),
Document(content="Haystack pipelines are directed multigraphs of components.", meta={"topic": "software"}),
])
print(store.count_documents())
3
InMemoryDocumentStore keeps documents in process memory. It is for prototyping, tutorials and tests; it disappears when the process exits and is not shared between server replicas. It does have save_to_disk and load_from_disk if you want a prototype to survive a restart. For anything real you swap in an integration package — Qdrant, pgvector, Weaviate, Chroma and many others are supported — and the rest of your pipeline does not change, which is the point of the store being an interface.
Notice the meta dictionary. You will use it constantly: to filter retrieval by language or department or date, to show the source of an answer, and to separate one tenant's documents from another's. Metadata filters in Haystack have a uniform shape. A comparison filter is a dictionary with three keys:
{"field": "meta.topic", "operator": "==", "value": "geography"}
The operators are ==, !=, >, >=, <, <=, in and not in. A logic filter combines them:
{"operator": "AND", "conditions": [
{"field": "meta.topic", "operator": "==", "value": "geography"},
{"field": "meta.year", "operator": ">=", "value": 2024},
]}
The logical operators are AND, OR and NOT. You can pass filters straight to store.filter_documents(filters=...) or to any retriever's filters argument. Support varies by store — Chroma, for example, has no NOT but adds a contains operator — and an operator a store does not support raises FilterError. Writing "gt" instead of ">" raises the same error, and that typo is common enough to be worth memorising.
Now the first pipeline. Keyword search, no model, no API key, nothing to pay for.
from haystack import Pipeline
from haystack.components.retrievers import InMemoryBM25Retriever
pipe = Pipeline()
pipe.add_component("retriever", InMemoryBM25Retriever(document_store=store, top_k=2))
result = pipe.run({"retriever": {"query": "Which city is on the Nile?"}})
for doc in result["retriever"]["documents"]:
print(round(doc.score, 3), doc.content)
3.147 Cairo sits on the Nile and is the capital of Egypt.
1.882 Riyadh is the capital of Saudi Arabia.
Your scores will differ; BM25 scores are not normalised and depend on the corpus. What matters is the shape of the call. pipe.run() takes a dictionary keyed by component name, whose values are dictionaries of that component's inputs. The result comes back keyed the same way, then by output socket name. result["retriever"]["documents"] is a list of Document objects with .score filled in.
BM25 is a keyword-ranking algorithm: it scores documents by how often the query's words appear in them, discounted by how common those words are across the corpus. It finds nothing when the user's wording does not overlap the document's. That is the limitation embeddings fix, later in this guide.
include_outputs_from={"retriever", "prompt_builder"} to run(). This one argument is the most useful debugging tool in the library.
- Build the store and the search pipeline above.
- Add
filters={"field": "meta.topic", "operator": "==", "value": "software"}to the retriever's run inputs and search again. - Change the operator to
"gt"and read the exception.
FilterError, and recognising that class by name is worth more than remembering the operator list.
Prompts, chat messages and your first RAG pipeline
Three components, two connections, and you have the application everyone wants to build.
from haystack import Pipeline
from haystack.components.retrievers import InMemoryBM25Retriever
from haystack.components.builders import ChatPromptBuilder
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.dataclasses import ChatMessage
from haystack.utils import Secret
template = [ChatMessage.from_user(
"Answer the question using only these documents. If they do not contain the "
"answer, say you do not know.\n\n"
"{% for doc in documents %}- {{ doc.content }}\n{% endfor %}\n"
"Question: {{ question }}"
)]
rag = Pipeline()
rag.add_components({
"retriever": InMemoryBM25Retriever(document_store=store, top_k=3),
"prompt_builder": ChatPromptBuilder(template=template, required_variables="*"),
"llm": OpenAIChatGenerator(
api_key=Secret.from_env_var("OPENAI_API_KEY"),
model="gpt-5-mini",
),
})
rag.connect_many([
("retriever", "prompt_builder.documents"),
("prompt_builder", "llm"),
])
question = "Which city is on the Nile?"
result = rag.run({
"retriever": {"query": question},
"prompt_builder": {"question": question},
})
print(result["llm"]["replies"][0].text)
Cairo sits on the Nile and is the capital of Egypt.
Several things in that listing deserve attention.
add_components({...}) and connect_many([...]) are new in Haystack 3.2. They are convenience methods: add_components is atomic, so if any one component fails a check nothing is added, and connect_many connects in order and stops at the first failure, keeping the connections it already made. add_component now returns the pipeline, so you can chain calls. On any version before 3.2 you write add_component and connect one at a time, which is exactly the same graph in more lines.
The connection strings use "sender.output" and "receiver.input" syntax. ("retriever", "prompt_builder.documents") leaves the sender's socket unnamed because the retriever has only one output, so there is no ambiguity. ("prompt_builder", "llm") names neither side for the same reason. When more than one pairing is possible Haystack refuses and tells you to be explicit, which is far better than guessing.
The template is a list[ChatMessage] containing Jinja. {% for doc in documents %} loops over the documents the retriever sent; {{ question }} is a variable you supply at run time. The two template variables map to two inputs of prompt_builder, which is why the question appears twice in the run() dictionary — once as the retriever's search query, once as the prompt's variable. They are genuinely two different inputs that happen to hold the same string, and seeing that clearly is a good sign you have understood the graph.
ChatPromptBuilder, PromptBuilder, Agent and LLM all default to required_variables="*". Omit one and you get ValueError: Missing mandatory input '<var>' for component '<name>'. In 2.x every variable was optional, so a typo'd variable name silently rendered an empty string and the model answered from nothing. The new behaviour is a real improvement. If you need the old one, pass required_variables=None; to require only some, pass a list of names.
The generator is worth a paragraph of its own. OpenAIChatGenerator's default model is gpt-5-mini, so the explicit model= above is documentation rather than necessity. Its api_key takes a Secret, and Secret.from_env_var("OPENAI_API_KEY") is what you should almost always use: it stores only the variable name, so a serialised pipeline contains no credential. Its sibling Secret.from_token("sk-...") holds the literal value and cannot be serialised at all — attempting it raises Cannot serialize token-based secret. Use an alternative secret type like environment variables. Treat from_token as a thing for throwaway notebooks.
api_base_url is the other argument worth knowing on day one. Point it at any OpenAI-compatible server — vLLM, Ollama, LM Studio, LiteLLM — and the same component talks to a local or self-hosted model. For teams in the Gulf and Egypt working under data-residency rules that forbid sending customer text to a model hosted abroad, this single argument is often the difference between a pilot that ships and one that stalls in legal review. Everything else in your pipeline stays identical.
The reply is a ChatMessage, not a string. .text gives the text, .meta gives metadata such as finish reason and token counts, and .tool_calls gives tool calls when the model requests them. Haystack 2.x returned a separate result["meta"] list alongside the replies; that list is gone, and the metadata now lives on each message where it belongs.
- Run the RAG pipeline. If you have no API key, replace the generator with
MockChatGenerator()fromhaystack.components.generators.chatand confirm the pipeline runs end to end. - Delete
"prompt_builder": {"question": question}from the run inputs and read the error. - Add
include_outputs_from={"prompt_builder"}and print the rendered prompt.
Indexing real files
Three hand-written documents are not a corpus. A real indexing pipeline converts files into documents, cleans them, cuts them into retrievable chunks and writes them to the store.
from pathlib import Path
from haystack import Pipeline
from haystack.components.converters import TextFileToDocument
from haystack.components.preprocessors import DocumentCleaner, DocumentSplitter
from haystack.components.writers import DocumentWriter
from haystack.document_stores.types import DuplicatePolicy
indexing = Pipeline()
indexing.add_components({
"converter": TextFileToDocument(),
"cleaner": DocumentCleaner(),
"splitter": DocumentSplitter(split_by="word", split_length=200, split_overlap=40),
"writer": DocumentWriter(document_store=store, policy=DuplicatePolicy.OVERWRITE),
})
indexing.connect_many([
("converter", "cleaner"),
("cleaner", "splitter"),
("splitter", "writer"),
])
files = list(Path("corpus").glob("*.txt"))
report = indexing.run({"converter": {"sources": files}})
print(report["writer"]["documents_written"])
184
Each stage earns its place. Converters turn files or byte streams into documents; besides TextFileToDocument there are PyPDFToDocument (needs pypdf), HTMLToDocument (needs trafilatura) and MarkdownToDocument. DocumentCleaner strips the debris that wrecks retrieval: repeated headers and footers, runs of whitespace, empty lines. Haystack 3.2 added min_content_length= so it can drop chunks too short to be useful.
DocumentSplitter is where the thinking happens. Retrieval works on chunks, and chunk size is the one knob with the largest effect on answer quality. split_by accepts word, sentence, line, page, passage, period, function, and — new in 3.2 — token, which needs pip install tiktoken and takes a tokenizer_encoding such as "o200k_base". Token splitting is the honest option when you care about fitting a model's context window, because that is the unit the model actually counts.
split_overlap repeats a tail of each chunk at the head of the next, so a sentence that straddles a boundary is not lost to both sides. Note the guard rail: split_overlap >= split_length raises ValueError at construction, not at run time, which is a small kindness. Richer splitters exist when the defaults are not enough — RecursiveDocumentSplitter, MarkdownHeaderSplitter, EmbeddingBasedDocumentSplitter — and DocumentPreprocessor bundles cleaning and splitting into one component.
DocumentWriter takes the store and a DuplicatePolicy. The default is NONE, which defers to whatever the store does. The others are OVERWRITE, SKIP and FAIL. Choose deliberately: re-running an indexing pipeline with OVERWRITE is idempotent and safe, while the default can leave you with two copies of everything and silently halved retrieval precision.
Document with no explicit id gets one hashed from its content and metadata. In 3.0 that hash began using canonical key-sorted JSON of meta, and non-JSON values pass through str() instead of repr(). The consequence is blunt: any document with non-empty metadata gets a different ID than it did in 2.x. Deduplication against data indexed under 2.x therefore fails silently — SKIP skips nothing, OVERWRITE writes a second copy. Re-ingest after upgrading, or pass explicit id= values you control.
- Make a
corpus/folder with two or three.txtfiles of a few hundred words each. - Run the indexing pipeline and note the number of documents written.
- Change
split_lengthto 50 and re-run, then to 500. - Ask the same question of your RAG pipeline after each run.
Embeddings and semantic retrieval
BM25 matches words. Ask "what is the capital of Egypt?" of a corpus that says "Cairo is Egypt's seat of government" and keyword search may well miss it. Embeddings fix this by mapping text to a vector such that similar meanings land near each other, and retrieval becomes a nearest-neighbour search.
This needs two components rather than one, and the distinction confuses everybody once. A document embedder takes documents and returns them with .embedding filled in; it runs at indexing time. A text embedder takes a query string and returns a bare embedding; it runs at query time. They must use the same model, or the two sets of vectors live in incompatible spaces and your retrieval is noise that looks plausible.
from haystack import Pipeline
from haystack.components.embedders import OpenAIDocumentEmbedder, OpenAITextEmbedder
from haystack.components.retrievers import InMemoryEmbeddingRetriever
from haystack.components.writers import DocumentWriter
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.document_stores.types import DuplicatePolicy
vector_store = InMemoryDocumentStore(embedding_similarity_function="cosine")
MODEL = "text-embedding-ada-002"
index = Pipeline()
index.add_components({
"embedder": OpenAIDocumentEmbedder(model=MODEL),
"writer": DocumentWriter(document_store=vector_store, policy=DuplicatePolicy.OVERWRITE),
})
index.connect("embedder", "writer")
index.run({"embedder": {"documents": my_documents}})
query = Pipeline()
query.add_components({
"text_embedder": OpenAITextEmbedder(model=MODEL),
"retriever": InMemoryEmbeddingRetriever(document_store=vector_store, top_k=3),
})
query.connect("text_embedder.embedding", "retriever.query_embedding")
hits = query.run({"text_embedder": {"text": "What is Egypt's seat of government?"}})
for doc in hits["retriever"]["documents"]:
print(round(doc.score, 3), doc.content[:60])
OpenAIDocumentEmbedder and OpenAITextEmbedder both default to the model text-embedding-ada-002. Set the model explicitly anyway, as above, in one constant used by both pipelines — that single MODEL variable is the cheapest possible insurance against the mismatch described a moment ago.
InMemoryDocumentStore has an embedding_similarity_function argument, "dot_product" by default with "cosine" as the alternative. Cosine ignores vector length and compares direction only, which is usually what you want for text. InMemoryEmbeddingRetriever takes top_k (10 by default), filters, scale_score and return_embedding, and its run input is query_embedding rather than query — that difference in socket name is precisely what the explicit connect("text_embedder.embedding", "retriever.query_embedding") is spelling out.
Embedding retrieval
- Finds paraphrases and synonyms
- Survives different vocabulary
- Needs an embedding model at index and query time
- Costs money or GPU time per document
- Weak on exact identifiers and rare codes
BM25 keyword retrieval
- Exact on names, codes and part numbers
- Free, instant, no model
- Nothing to index beyond the text
- Misses every paraphrase
- Collapses when wording differs
The column headings say "good" and "bad" because the styling needs two columns, not because one wins. Mature systems run both and combine the results. Building that is mid-level work; knowing that it is the normal destination keeps you from over-investing in either one now.
If you would rather not pay per embedding, the Sentence Transformers integration runs models locally: pip install sentence-transformers-haystack, then import from haystack_integrations.components.embedders.sentence_transformers. Note the import path carefully — these components lived in haystack in 2.x and moved out in 3.0. About thirty components made that move, and a stale import is the single most common 3.0 migration failure.
- Index five documents with the embedding pipeline.
- Ask a question that deliberately shares no words with any document.
- Ask the same question of a BM25 retriever over the same documents.
Reading Haystack's errors
Haystack's error messages are unusually informative, and learning to read four or five families of them turns most problems into two-minute fixes.
Connection errors. PipelineConnectError fires at connect() time, before anything runs. The messages are specific. Cannot connect 'a.out' with 'b.in': their declared input and output types do not match. means exactly what it says; fix the types or insert an OutputAdapter. more than one connection is possible between these components means you left both socket names off and the pipeline will not guess — name them. Component 'b' cannot accept multiple inputs to 'x'. It is already connected to component 'a', and it can only accept inputs from multiple senders if its type is list, Optional[list], or union of list types. is the rule about variadic sockets: only list-typed inputs can take several senders, and the pipeline then concatenates the lists. And 'a' does not have any output connections. Please check that the output types of 'a.run' are set, for example by using the '@component.output_types' decorator. is the missing-decorator mistake you provoked deliberately earlier.
Naming errors. ValueError: A component named 'x' already exists in this pipeline: choose another name. Names cannot contain a dot — it is the socket separator — and _debug is reserved. One more, easy to hit and hard to guess: Component instance cannot be added to the pipeline more than once. A single instance belongs to one pipeline under one name. If you want the same component twice, construct it twice.
Missing input errors. ValueError: Missing mandatory input '<var>' for component '<name>'. Either a template variable is required and absent, or a run parameter without a default was neither passed nor connected. Check both before changing required_variables.
Runtime errors. PipelineRuntimeError wraps whatever a component actually threw:
The following component failed to run:
Component name: 'llm'
Component type: 'OpenAIChatGenerator'
Error: ...
Read the inner error; the wrapper only tells you where. The exception also carries .pipeline_snapshot, so you can resume rather than restart — genuinely useful when the failure happened after an expensive embedding step. A related message, returned an invalid output ... Expected a dictionary, but got X instead, means one of your own components forgot to return a dict.
Authentication errors. ValueError: None of the following authentication environment variables are set: ('OPENAI_API_KEY',). A strict environment-variable Secret with nothing in the variable. Since 3.0 this appears at warm_up() or the first run rather than at construction, because clients are built at warm-up. Export the key, or use Secret.from_env_var(..., strict=False) for genuinely optional credentials.
Loop errors. PipelineMaxComponentRuns: Maximum run count 100 reached for component '<name>' means a loop never exited. Pipeline(max_runs_per_component=...) sets that limit, 100 by default, and lowering it while debugging gets you to the problem faster. Its cousin is PipelineComponentsBlockedError, whose message lists the suspects and says the cause is either no valid entry point or a circular dependency. In practice it is usually a required template variable that only a later component produces, so nothing can start.
Deserialisation errors. DeserializationError: Refusing to deserialize a class from module '<m>': the module is not on the trusted-module allowlist. That is the 3.0 security model, covered next. Its sibling, Refusing to deserialize unknown parameter '<key>' for '<Class>', is what stale YAML looks like now — a parameter that was removed from a class in a newer release will stop a load instead of being quietly ignored.
- Add the same retriever instance to a pipeline twice under two names.
- Name a component
my.retriever. - Connect the retriever's documents output directly to the generator's
messagesinput.
Configuration, secrets and saving a pipeline
A pipeline can be written out as YAML and read back. This is one of Haystack's distinguishing features and it is worth understanding properly, including the safety machinery around it.
yaml_text = rag.dumps()
with open("rag.yaml", "w") as fp:
rag.dump(fp)
The resulting file has a predictable structure: a components mapping of names to {type, init_parameters}, a connections list of {sender, receiver} pairs, plus max_runs_per_component, metadata and connection_type_validation. Secrets serialise by reference, never by value:
components:
llm:
type: haystack.components.generators.chat.openai.OpenAIChatGenerator
init_parameters:
model: gpt-5-mini
api_key:
type: env_var
env_vars: [OPENAI_API_KEY]
strict: true
connections:
- sender: retriever.documents
receiver: prompt_builder.documents
Only the variable name is stored. That file is safe to commit, which is the entire reason to prefer Secret.from_env_var over Secret.from_token.
Loading is where Haystack 3.0 changed behaviour deliberately:
from haystack import Pipeline
with open("rag.yaml") as fp:
loaded = Pipeline.load(fp)
Loading a pipeline is gated by a trusted-module allowlist. The defaults are haystack, haystack_integrations, haystack_experimental, builtins, typing and collections. Dangerous builtins — eval, exec, compile, __import__, open, getattr — are blocked outright. And nested init_parameters keys are validated against the parent class's constructor signature, so a stale or misspelled key fails loudly.
The reason is straightforward once stated: a pipeline file names Python classes to import and arguments to pass them, so a YAML file is executable code. Before 3.0, loading an untrusted pipeline file was equivalent to running an untrusted script. If your own components live outside the allowlist, extend it explicitly with allowed_modules=["mypkg.*"], or set HAYSTACK_DESERIALIZATION_ALLOWLIST="mypkg.*". There is a global escape hatch, unsafe=True, and you should treat it as off-limits for any file you did not write yourself.
Default serialisation has one requirement that catches people: constructor parameter names must match instance attribute names, one to one. A component whose __init__ takes model_name but stores self.model cannot be serialised by the default machinery. Assign self.<param> = <param> for every argument and the problem disappears.
The environment variables worth knowing on day one:
| Variable | Effect |
|---|---|
OPENAI_API_KEY |
Read by the OpenAI components through Secret |
OPENAI_TIMEOUT |
Request timeout, default 30 seconds |
OPENAI_MAX_RETRIES |
Retry count, default 5 |
HAYSTACK_TELEMETRY_ENABLED=False |
Opt out of anonymous usage telemetry |
HAYSTACK_LOGGING_USE_JSON=true |
Force JSON logs; automatic when there is no TTY |
HAYSTACK_DESERIALIZATION_ALLOWLIST |
Extend the trusted-module list for loading |
On telemetry: Haystack sends anonymous component-usage events to PostHog on an EU host, identified by a random UUID, with no IP address, hostname, file paths, queries or content. Many employers in the region will still want it off, and one environment variable does it. The configuration file lives at ~/.haystack/config.yaml.
Logging deserves a note because 3.0 changed it for the better. Importing Haystack no longer reconfigures the root logger or global structlog, so the library can no longer silently reformat your application's logs. The default level is WARNING. To see more:
import logging
logging.getLogger("haystack").setLevel(logging.DEBUG)
If you actually wanted the old global behaviour, configure_logging(logger_name="") restores it, and configure_logging(propagate=False) is the fix for duplicated log lines.
- Dump your RAG pipeline to
rag.yamland read the file. - Confirm your API key does not appear anywhere in it.
- Add a nonsense key such as
temperture: 0.5under the generator'sinit_parametersand load the file. - Rebuild the generator with
Secret.from_token("sk-test")and calldumps().
DeserializationError naming the unknown parameter — the typo-catching behaviour that 2.x lacked. Step four refuses with the token-secret message. Both failures are the library protecting you.
Seeing what the pipeline did
Two tools: a picture of the graph, and a trace of the run.
pipe.show() # renders inline in Jupyter
pipe.draw(path=Path("pipeline.png")) # writes a PNG
from pathlib import Path
rag.draw(path=Path("rag.png"))
draw and show take keyword-only arguments in 3.2. You will find documentation examples written as draw("my_pipeline.png", ...); that positional form fails against the current signature. Write draw(path=...) and move on.
Both methods render the graph by sending it to mermaid.ink, a public service on the internet. That is a privacy problem for a pipeline whose component names and configuration you would rather not publish, and it is also why draw() fails on a machine with no network access. Run the renderer yourself:
docker run --platform linux/amd64 --publish 3000:3000 --cap-add=SYS_ADMIN \
ghcr.io/jihchi/mermaid.ink
rag.draw(path=Path("rag.png"), server_url="http://localhost:3000")
draw accepts mermaid-image (the default) or mermaid-text, and the text form needs no server at all — useful in CI, where you can diff the graph as text.
For understanding a run rather than a structure, start with include_outputs_from:
result = rag.run(
{"retriever": {"query": q}, "prompt_builder": {"question": q}},
include_outputs_from={"retriever", "prompt_builder"},
)
print(result["prompt_builder"]["prompt"][0].text)
That prints the exact prompt sent to the model, with the retrieved passages interpolated. Get into the habit of looking at it whenever an answer is wrong, because the most common cause of a bad answer is a prompt full of the wrong documents, and no amount of prompt engineering fixes a retrieval problem.
pipeline.inputs() is the other small tool worth knowing: it lists what the pipeline expects and which of those inputs are mandatory. When you cannot work out what key run() wants, ask.
For anything beyond this — span-level traces, token accounting, prompt history across sessions — you want a real observability tool, and Haystack integrates with several. Langfuse is the common choice in this series; note that it needs HAYSTACK_CONTENT_TRACING_ENABLED=true, set before importing Haystack, because recording inputs and outputs is off by default for the good reason that spans would otherwise contain prompts and personal data. Tracing is never enabled automatically in 3.x; if you remember HAYSTACK_AUTO_TRACE_ENABLED from 2.x, it was removed and now does nothing.
- Draw your RAG pipeline and look at the picture.
- Run it with
include_outputs_fromcovering every component and print the rendered prompt. - Call
rag.inputs()and compare the result to the dictionary you pass torun().
Putting it all together
One script, two pipelines, one store: index a folder of text files and answer questions about them with citations. This is the thing to keep and modify.
import logging
from pathlib import Path
from haystack import Pipeline
from haystack.components.builders import ChatPromptBuilder
from haystack.components.converters import TextFileToDocument
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.components.preprocessors import DocumentCleaner, DocumentSplitter
from haystack.components.retrievers import InMemoryBM25Retriever
from haystack.components.writers import DocumentWriter
from haystack.dataclasses import ChatMessage
from haystack.document_stores.in_memory import InMemoryDocumentStore
from haystack.document_stores.types import DuplicatePolicy
from haystack.utils import Secret
logging.getLogger("haystack").setLevel(logging.INFO)
store = InMemoryDocumentStore()
indexing = Pipeline()
indexing.add_components({
"converter": TextFileToDocument(),
"cleaner": DocumentCleaner(),
"splitter": DocumentSplitter(split_by="word", split_length=180, split_overlap=30),
"writer": DocumentWriter(document_store=store, policy=DuplicatePolicy.OVERWRITE),
})
indexing.connect_many([
("converter", "cleaner"),
("cleaner", "splitter"),
("splitter", "writer"),
])
template = [
ChatMessage.from_system(
"You answer strictly from the supplied documents. "
"If they do not contain the answer, say so plainly."
),
ChatMessage.from_user(
"{% for doc in documents %}"
"[{{ loop.index }}] {{ doc.meta.file_path }}\n{{ doc.content }}\n\n"
"{% endfor %}"
"Question: {{ question }}\n"
"Cite the sources you used by their bracketed number."
),
]
answering = Pipeline()
answering.add_components({
"retriever": InMemoryBM25Retriever(document_store=store, top_k=4),
"prompt_builder": ChatPromptBuilder(template=template, required_variables="*"),
"llm": OpenAIChatGenerator(
api_key=Secret.from_env_var("OPENAI_API_KEY"),
model="gpt-5-mini",
),
})
answering.connect_many([
("retriever", "prompt_builder.documents"),
("prompt_builder", "llm"),
])
def ask(question: str, show_prompt: bool = False) -> str:
extras = {"prompt_builder"} if show_prompt else set()
result = answering.run(
{"retriever": {"query": question}, "prompt_builder": {"question": question}},
include_outputs_from=extras,
)
if show_prompt:
print("--- prompt ---")
print(result["prompt_builder"]["prompt"][-1].text)
print("--- /prompt ---")
return result["llm"]["replies"][0].text
if __name__ == "__main__":
files = sorted(Path("corpus").glob("*.txt"))
if not files:
raise SystemExit("Put some .txt files in ./corpus first.")
report = indexing.run({"converter": {"sources": files}})
print(f"indexed {report['writer']['documents_written']} chunks from {len(files)} files")
answering.warm_up()
answering.dump(open("answering.yaml", "w"))
while True:
try:
question = input("\nquestion> ").strip()
except (EOFError, KeyboardInterrupt):
break
if not question:
break
print(ask(question, show_prompt=question.startswith("?")))
answering.close()
Walk through the deliberate choices, because each one is a habit worth keeping.
The two pipelines share store and nothing else. Indexing runs once at startup; in a real deployment it would be a separate job run when the corpus changes. DuplicatePolicy.OVERWRITE makes re-running it safe.
The template is two messages, not one. The system message carries the instruction that must not be negotiable — answer only from the documents — and the user message carries the data. Separating them is the right default: instructions that live in the same turn as retrieved text are easier for that text to override.
The prompt includes doc.meta.file_path and a loop.index number, and asks the model to cite the brackets. That metadata is there because TextFileToDocument put it there, and it travels through cleaning and splitting untouched. Citations are the cheapest trust mechanism a RAG system has, and they cost one line of Jinja.
answering.warm_up() is called explicitly before the loop rather than being left to the first run. In a server you do this at startup so the first user request is not the one that pays for client construction and model loading. answering.close() releases HTTP clients on the way out. Both methods exist on Pipeline as well as on components.
show_prompt is wired to a leading ? on the question. Build the debug path in from the start; you will use it on the first bad answer, and a flag you already have beats an edit under pressure.
And the pipeline is dumped to answering.yaml on every run, which gives you a reviewable artefact of exactly what shipped — a file a colleague can read without reading your Python.
- Run
app.pyagainst a folder of your own notes. - Ask a question whose answer is genuinely absent and check whether the model admits it.
- Prefix a question with
?and read the prompt. - Swap
InMemoryBM25Retrieverfor the embedding retriever from the semantic section and compare the answers.
What you can now do, and what comes next
You can install Haystack and confirm its version. You can write a component with declared input and output sockets and test it without a pipeline. You can build a pipeline, connect named sockets, read the connection errors when the types do not fit, and get at the intermediate outputs that run() hides by default. You can index a folder of files through a converter, cleaner, splitter and writer, with a duplicate policy you chose on purpose. You can retrieve by keyword and by embedding, filter by metadata, and explain to a colleague when each kind of retrieval wins. You can build a chat prompt from a template, call a model through a Secret that never lands in a file, and read the reply off a ChatMessage. You can serialise the whole graph to YAML and load it back, and you know why loading is gated by an allowlist. You can read the six or seven error families Haystack actually produces. And you know which of your 2.x instincts are now wrong.
What you have not touched is the other half of the library. Agents are a loop — call the model, run the tools it asked for, check whether to exit — and Agent is the component that implements it, with max_agent_steps as a safety limit and an exit_reason output explaining how it stopped. Tools are functions the model can call, created with the @tool decorator or by wrapping a component with ComponentTool. State is the agent's shared scratchpad. Hooks let you intervene at defined points in that loop, including the human-in-the-loop confirmation you want before a tool sends an email or writes to a database. Haystack's Agent is deliberately not a graph framework; if you have met LangGraph, the trade-off between an explicit state machine and a bounded tool loop is worth thinking about once you have used both.
The production side is also ahead of you. Pipeline.run_async and Pipeline.stream replace the AsyncPipeline class that 3.0 removed, and they are what you want inside a FastAPI endpoint. Hayhooks is a separate package that serves pipelines and agents as REST, OpenAI-compatible, MCP and A2A endpoints, which is how a Haystack pipeline becomes something a front end can talk to. MCP works in both directions: mcp-haystack lets your agent consume MCP tools, and hayhooks mcp run exposes your pipelines as MCP tools for someone else's agent. For judging whether a change made answers better rather than merely different, Haystack ships evaluators and integrates with RAGAS.
Before any of that, do one thing: take app.py, point it at a corpus you actually care about, and use it for a week. Every concept in the mid-level guide — hybrid retrieval, reranking, async serving, loops and routers, agents with tools — is an answer to a problem you will meet in that week. Meeting the problem first is what makes the answer stick.
Sources
- Haystack documentation: Intro
- Installation
- Components
- Custom components
- Pipelines
- Creating pipelines
- Smart pipeline connections
- Document store
- Metadata filtering
- Secret management
- Serialization
- Visualizing pipelines
- Debugging pipelines
- Logging
- Tracing
- Telemetry
- Breaking change policy
- Migration guide
- FAQ
- Docker
- Haystack 3.2.0 release notes
- Haystack 3.0.0 release notes
- MIGRATION.md
- haystack-ai on PyPI
- deepset/haystack on Docker Hub