This is part one of three. It takes you from never having touched Pinecone to a working semantic search that you built, queried, filtered, and cleaned up yourself. By the end you can create an index, load records into it, search it by meaning, narrow results with filters, read the common error messages, and explain to a colleague why a search you just wrote returned nothing for the first few seconds. Mid-level and Senior take the same topics further; nothing here is thrown away.
Each section ends with a Try it task. Do them as you go. Vector databases feel abstract until you have watched your own sentence find a similar sentence, and that moment takes only a few minutes of work.
This guide was checked against the Pinecone API version 2026-07 and the Python SDK 10.0, the current releases at the time of writing. Pinecone moves fast: it ships a new stable API version every quarter, and the SDK major version follows it. Where something changed recently and old tutorials still show the old form, this guide says so and shows the current one.
What Pinecone is, and the problem it solves
Pinecone is a managed database for searching by meaning. You give it pieces of content (a paragraph, a product description, a support ticket) and it stores them in a form that lets you later ask, "which stored items are most like this one?" You do not run servers, tune memory, or patch anything. You call an API and Pinecone does the rest.
To see why that is useful, start with what came before. A classic database or search box matches words. If your help centre has an article titled "How to reset your password" and a user searches "I can't log in", a word-matching system finds nothing, because the two texts share no important words. A person sees at once that they are about the same thing. Teaching a computer that is the problem semantic search solves.
The trick is the embedding. An embedding model is a neural network that turns a piece of text (or an image, or audio) into a list of numbers, called a vector. The model is trained so that texts with similar meaning end up with similar lists of numbers. "Reset my password" and "I can't log in" land close together; "quarterly sales report" lands far away. Once everything is a list of numbers, "similar in meaning" becomes "close in space", and closeness is something a computer can calculate quickly.
The diagram is the whole idea. Everything else in this guide is detail about the third box.
Finding the closest vectors out of a few hundred is trivial; you can compare against all of them. Finding the closest out of a hundred million, in a few dozen milliseconds, is hard. Specialised data structures do it by being approximate: they find the nearest neighbours very reliably, but without checking every single vector. Pinecone's job is to build and run those structures for you, keep them updated as you add data, and keep them fast as you grow.
Two consequences of that design explain a lot of what follows. Notice them now.
Pinecone stores and searches vectors; it does not create meaning by itself. Something has to produce the vectors. You can bring your own from any embedding model, or you can let Pinecone host the embedding model and call it for you. Both paths are shown below.
The search is approximate and the database is eventually consistent. An approximate search can occasionally miss a neighbour that an exhaustive search would have found. And a record you just wrote may take a short while before it appears in search results. Neither is a bug. Both are trade-offs that buy speed and scale, and knowing about them saves hours of confusion.
What people use it for:
Retrieval for LLM apps
Find the passages most relevant to a question, then give them to a language model so it answers from your documents. This pattern is called retrieval-augmented generation, or RAG.
Semantic search
A search box that understands "cheap flights to Cairo" and "low-cost tickets to Egypt" as the same request.
Recommendations
Show products, articles or videos similar to the one the user is looking at.
Deduplication and clustering
Spot near-duplicate tickets or documents even when the wording differs.
Pinecone is one of several vector databases. Others in this catalogue include Qdrant, Weaviate, Milvus and Chroma, plus the library FAISS and the Postgres extension pgvector. Pinecone's particular position is that it is fully managed and serverless: there is nothing to deploy. That is a real advantage when your team is small and a real limitation when your data must stay inside infrastructure you control. We return to that trade-off near the end.
- Write five short sentences about three different topics, for example two about cooking, two about football and one about taxes.
- Pick one cooking sentence and rank the other four by how similar in meaning they are to it. Do it by gut feeling.
- Keep your ranking. You will compare it with Pinecone's ranking later in this guide.
The core mental model: five nouns
Pinecone has a small vocabulary. Learn these five nouns and the documentation becomes readable.
An organization is the account at the top. It groups projects that share billing, and it holds the people and the roles. A project belongs to exactly one organization and contains your indexes. API keys are scoped to a project, which means a key created in one project cannot see indexes in another. Beginners often create a key, then cannot find their index, because the index lives in a different project from the key's project.
An index is the container that actually holds your data. When you create one you choose a cloud (AWS, GCP or Azure) and a region, and that choice cannot be changed later. If you later need the data somewhere else, you create a new index and load it again. For a vector index you also choose the dimension, the length of the vectors it stores, and the metric, the way similarity is measured. Both are fixed at creation too.
A namespace is a partition inside an index. Every read and every write targets exactly one namespace, and a search only looks inside the namespace you name. You do not have to create namespaces in advance: the first time you write into a name that does not exist, Pinecone creates it. Namespaces are how you keep customers apart (one namespace per customer) or keep environments apart (one for staging data, one for production data) without paying for many indexes. If you do not name a namespace, you are using the default one, whose name is the empty string, which newer API versions display as __default__.
A record is one stored item. It has an id (a string you choose, up to 512 ASCII characters), the values (the dense vector), and optionally metadata, a small JSON object of facts about the item such as its category, year or source URL. Metadata does two jobs: you can filter by it at search time, and you can return it with the results so your application knows what the match was.
A few more words come up everywhere.
- Dense vector: the ordinary kind of embedding, a fixed-length list of decimal numbers such as 1,024 of them. It captures meaning.
- Sparse vector: a list of mostly zeros, stored as pairs of position and weight. It captures which specific words appear, and is used for keyword-style matching.
- Similarity metric:
cosine,dotproductoreuclidean. For text embeddings,cosineis the usual choice, and the embedding model's documentation tells you which one it was trained for. Witheuclideanthe returned score is a squared distance, so lower is better; with the other two, higher is better. - Top-k: how many nearest matches you ask for.
top_k=5means "give me the five closest". - Upsert: "update or insert". If a record with that ID exists, it is replaced; if not, it is created. There is no separate "insert" call.
On the current API version there are three kinds of index, chosen when you create it, and each kind speaks to its own set of data calls.
| Kind | You create it with | You store | You talk to it with |
|---|---|---|---|
| Vector index (the classic kind) | a dimension, a metric | records with vectors you computed | upsert, query, fetch, delete |
| Integrated-embedding index | an embedding model name | records with plain text, which Pinecone embeds | upsert_records, search |
Document index (new in 2026-07) |
a schema | JSON documents | index.documents.upsert, index.documents.search |
The kind is fixed for the life of the index, and a request of the wrong style is refused with an error that names the right one. You cannot convert an existing index from one kind to another; you create a new index and load the data again. This guide's first project uses the integrated kind because it lets you start with plain text and no embedding model of your own. Later sections show the classic and document kinds so you recognise them in other people's code.
index.query(...) with a vector, or create a classic index and call index.documents.search(...), you get an error from the API that names the right endpoint. Read the message; it tells you what to call instead.- Draw the five-noun chain on paper and write, under each noun, one real example from an app you might build, for example Project: "support-bot", Namespace: "customer-42".
- Decide which of your five sentences from the previous task would be the id, which the text, and which facts would be metadata.
Installing the tools and checking the setup
You need three things: an account with an API key, a client library, and optionally the command-line tool.
The account and the key. Sign up at pinecone.io. The Starter plan is free and needs no card, which is plenty for everything in this guide. In the console, open your project and choose API keys, then create a key. The key is shown once. Copy it immediately and store it somewhere safe. If you lose it, you delete it and create another.
Treat the key like a password. Never paste it into source code that you commit, and never put it in JavaScript that runs in a browser. Put it in an environment variable instead. On macOS or Linux:
export PINECONE_API_KEY="paste-your-key-here"
On Windows PowerShell the same thing is:
$env:PINECONE_API_KEY="paste-your-key-here"
The variable lasts only for that terminal window. To make it permanent, add the line to your shell profile or use your operating system's environment settings, and make sure no file containing the key is tracked by Git.
The Python client. The official Python package is called pinecone. It needs Python 3.10 or newer. Create a clean virtual environment so the install does not disturb anything else:
python -m venv .venv
source .venv/bin/activate # Windows PowerShell: .venv\Scripts\Activate.ps1
pip install --upgrade pinecone
Verify it:
python -c "import pinecone; print(pinecone.__version__)"
This guide targets version 10. If you see a lower major version, upgrade; the code below will not match older releases. Some tutorials mention pip install "pinecone[grpc]" or the old package name pinecone-client. Ignore them. The old name stopped receiving updates, and from version 9 onward the faster gRPC transport is built into the normal package, so no extras are needed. Pip may even warn you about unknown extras if you ask for them.
Now check that the library can reach Pinecone with your key:
from pinecone import Pinecone
pc = Pinecone() # reads PINECONE_API_KEY from the environment
print(pc.list_indexes()) # an empty listing on a fresh project
On a new project this prints an empty result and that is a success: it proves the key is valid and the network path works. If the variable is missing, the constructor stops with PineconeValueError: No API key provided. Pass api_key='...' or set the PINECONE_API_KEY environment variable. Export the variable in the same terminal you run Python from; a variable set in another window is invisible here.
The command-line tool (optional). Pinecone has a CLI called pc. On macOS or Linux with Homebrew:
brew tap pinecone-io/tap
brew install pinecone-io/tap/pinecone
Or use the install script from the official site:
curl -fsSL https://pinecone.io/install.sh | sh
On Windows, download the prebuilt binary from the releases page of the pinecone-io/cli repository on GitHub and put it on your PATH. Then check it and sign in:
pc version
pc auth login
pc auth status
pc index list
pc auth login opens your browser. The CLI is handy for inspecting what your code created, but everything in this guide can be done from Python, so skip it if you prefer.
export PINECONE_API_KEY=... in a file called .env.local that you add to .gitignore, and run source .env.local at the start of each session. The key never touches a file that Git tracks.- Create a free account, make an API key, and store it in an environment variable.
- Install the Python package in a fresh virtual environment and print its version.
- Run the three-line script that lists your indexes and confirm it prints without an error.
Your first index and your first search
Now the real thing. The goal is a small semantic search over six sentences, created from scratch. We use an integrated-embedding index, so Pinecone embeds the text for us and you never handle a vector.
Create a file called first_search.py. Start with the index:
from pinecone import Pinecone
pc = Pinecone()
index_name = "quickstart-demo"
if not pc.has_index(index_name):
pc.create_index_for_model(
name=index_name,
cloud="aws",
region="us-east-1",
embed={
"model": "llama-text-embed-v2",
"field_map": {"text": "chunk_text"},
},
)
index = pc.index(index_name)
Read it slowly, because every line carries an idea.
pc.has_index checks whether the name is taken, so running the script twice does not crash on an "already exists" error. create_index_for_model creates an index tied to a hosted embedding model. The model is llama-text-embed-v2, one of the models Pinecone runs for you. The field_map says which field in your records holds the text to embed: here, a field named chunk_text. The cloud and region (aws and us-east-1) are fixed forever. On the free Starter plan us-east-1 is the only region available, so that is what we use.
In the current SDK, creating an index waits until the index is ready before it returns, so the next line can use it straight away. Older tutorials add a polling loop; you do not need it any more. The last line, pc.index(index_name), gives you a handle for reading and writing data. In production you point at an index by its host address instead of its name, which saves a lookup, but the name is fine while you learn.
Index names follow rules: lowercase letters, digits and hyphens only, at most 45 characters, and they cannot start or end with a hyphen. No dots, no underscores, no capitals.
Now add data. Each record carries an ID, the text in the field we mapped, and some metadata. Notice that there are no vectors anywhere:
records = [
{"_id": "r1", "chunk_text": "To reset your password, open Settings and choose Security.", "topic": "account"},
{"_id": "r2", "chunk_text": "Refunds are issued to the original payment method within five days.", "topic": "billing"},
{"_id": "r3", "chunk_text": "Our mobile app supports Arabic and English interfaces.", "topic": "product"},
{"_id": "r4", "chunk_text": "You can change the email address linked to your account in Profile.", "topic": "account"},
{"_id": "r5", "chunk_text": "Invoices are emailed on the first day of every month.", "topic": "billing"},
{"_id": "r6", "chunk_text": "Dark mode can be switched on from the Appearance menu.", "topic": "product"},
]
index.upsert_records("demo", records)
upsert_records("demo", records) writes the records into the namespace called demo, creating it if needed. In records for this kind of index, the ID goes in the _id field and everything else is either the text field you mapped or metadata. Pinecone embeds each chunk_text with the hosted model and stores the resulting vectors next to the metadata.
Here is the step that surprises almost every beginner. If you search immediately after the upsert, you may see no results at all. Pinecone acknowledges a write as soon as it is safely recorded, but making it searchable takes a moment, because the system is eventually consistent. Wait a few seconds before searching:
import time
time.sleep(10)
results = index.search(
namespace="demo",
query={"inputs": {"text": "I can't log in to my account"}, "top_k": 3},
fields=["chunk_text", "topic"],
)
for hit in results['result']['hits']:
print(hit["_id"], round(hit["_score"], 3), hit["fields"]["chunk_text"])
Run it with python first_search.py. You should see the password and email-address records near the top, even though your query contained neither "password" nor "reset". Exact scores depend on the hosted model, so read the ordering rather than the numbers: higher is a closer match for a cosine index. Compare the order with the ranking you made by gut in the previous section. Where they agree, the embedding model understood the sentences the way you did.
The query object has two parts. inputs holds the search text, which Pinecone embeds with the same model that embedded your records; that sameness is essential, because vectors from different models cannot be compared. top_k is how many hits you want. fields lists which stored fields to return; fields you do not list are left out, which keeps responses small.
time.sleep is fine for a tutorial and wrong for real code. Under load, freshness can take longer. Mid-level covers how to check that a write is visible; for now just know that "nothing found right after writing" almost always means "not yet searchable", not "lost".When you are finished experimenting, delete what you created, because the free plan allows only a handful of indexes:
pc.delete_index("quickstart-demo")
Deletion is immediate and cannot be undone. In a real project you protect important indexes with deletion protection, which makes the delete call fail until you switch it off on purpose.
- Run the script end to end and read the three hits it prints.
- Change the query to "how do I get my money back" and predict which record comes first before running it.
- Add two more records of your own and search again. Notice the sleep: remove it and see what happens on the first run.
- Delete the index when you are done.
Bringing your own vectors: the classic index
The integrated index is the shortest path, but most existing code, and most tutorials you will meet, use the classic vector index, where you compute the embeddings yourself with whatever model you like (a Hugging Face model, an API from an AI provider, or your own). It is worth learning because it shows precisely what Pinecone stores.
Creating one requires the two decisions we mentioned: dimension and metric. The dimension must equal the length of the vectors your model produces, and it cannot be changed later. A model that outputs 1,536 numbers needs a 1,536-dimension index. Pinecone allows up to 20,000. For a learning exercise we use tiny three-number vectors, so you can see every value:
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone()
name = "classic-demo"
if not pc.has_index(name):
pc.create_index(
name=name,
dimension=3,
metric="cosine",
spec=ServerlessSpec(cloud="aws", region="us-east-1"),
)
index = pc.index(name)
On the newest SDK the dimension, metric and spec arguments are treated as a convenience form that still works; the SDK's native form describes the index with a schema, which the next section shows. Both create the same kind of index. You will meet the spec=ServerlessSpec(...) form in a huge amount of existing code, so it is good to recognise it. The word serverless means you do not choose machine sizes: capacity grows and shrinks with your data and traffic, and you pay for what you store and use.
Now write three records. Each has an ID, a vector of exactly three numbers, and metadata:
index.upsert(
namespace="shapes",
vectors=[
{"id": "a", "values": [1.0, 0.0, 0.0], "metadata": {"genre": "red"}},
{"id": "b", "values": [0.9, 0.1, 0.0], "metadata": {"genre": "red"}},
{"id": "c", "values": [0.0, 0.0, 1.0], "metadata": {"genre": "blue"}},
],
)
And ask for the neighbours of a vector that points mostly in the first direction:
import time
time.sleep(10)
res = index.query(
namespace="shapes",
vector=[0.95, 0.05, 0.0],
top_k=2,
include_metadata=True,
include_values=False,
)
for match in res["matches"]:
print(match["id"], match["score"], match["metadata"])
You should get a and b, both with scores very close to 1.0 (cosine similarity of 1.0 means identical direction), and not c, which points somewhere else entirely. That is the entire mechanism of semantic search, scaled down. A real embedding model simply produces longer vectors in which each direction encodes some aspect of meaning.
Notice the arguments to query. include_values=False tells Pinecone not to send the vectors back; you almost never need them and they make the response heavy. include_metadata=True returns the facts you stored. The limits worth remembering: top_k can be at most 10,000, a response can be at most 4 MB, and a single upsert request is capped at 1,000 records or 2 MB, whichever you hit first. Big loads are therefore sent in batches.
400 error describing the mismatch. Check len(vector) against the index dimension, and make sure the same model produces both document and query vectors.- Create the three-dimension index and upsert the three records.
- Query with
[0.0, 0.1, 0.9]and predict the first match before you run it. - Change the metric to
euclideanon a new index and notice that scores are now squared distances where smaller is better. - Delete both demo indexes when finished.
The document index: the newest style
With API version 2026-07, Pinecone introduced a third kind of index, the document index, and its Documents API. Instead of a record with a vector and a metadata blob, you store a JSON document: an object with a required _id, fields that Pinecone ranks on, and any other fields, which are stored and indexed automatically as filterable metadata. The same index can hold a dense vector, a sparse vector and text fields that support keyword search (BM25, the classic ranking formula used by search engines), all described by a schema you declare when you create the index.
Here is the beginner path in Python:
from pinecone import Pinecone, SchemaBuilder, DenseVectorQuery
pc = Pinecone()
schema = (
SchemaBuilder()
.add_dense_vector_field("embedding", dimension=3, metric="cosine")
.add_string_field("body", full_text_search={"language": "en"})
.build()
)
pc.indexes.create(
name="articles",
schema=schema,
deployment={"deployment_type": "managed", "cloud": "aws", "region": "us-east-1"},
)
index = pc.index("articles")
index.documents.upsert(
namespace="ns1",
documents=[
{"_id": "doc1", "embedding": [1.0, 0.0, 0.0], "body": "Intro to vector search", "year": 2024},
{"_id": "doc2", "embedding": [0.0, 1.0, 0.0], "body": "Cooking with lentils", "year": 2022},
],
)
The schema is the new idea. add_dense_vector_field declares a field named embedding holding vectors of three numbers. add_string_field with full_text_search declares a text field that supports keyword search. The deployment says where the index lives; managed is the serverless kind. The field year is not in the schema, yet it works as a filterable field automatically, because every undeclared field is stored as metadata. That is a major simplification compared with the classic style.
Searching combines a scoring method and optional filters:
import time
time.sleep(10)
by_meaning = index.documents.search(
namespace="ns1",
top_k=2,
score_by=[DenseVectorQuery(field="embedding", values=[0.9, 0.1, 0.0])],
filter={"year": {"$gte": 2024}},
include_fields=["body"],
)
for m in by_meaning.matches:
print(m.id, m.score)
by_words = index.documents.search(
namespace="ns1",
top_k=2,
score_by=[{"type": "text", "fields": ["body"], "query": "vector search"}],
include_fields=["body"],
)
Each document search uses one scoring method: dense vector, sparse vector, text, or a query-string. By default, only _id and _score come back; include_fields=["*"] returns every stored field. If you want meaning and keywords combined, you run two searches and merge the lists in your own code; that is a topic for the Mid-level guide.
Know the limits before you commit to this style. The schema cannot be changed after creation; it declares ranking fields only, and declaring a metadata-only field is rejected with a 400. A document index has no conversion path from an older index, so adopting it means creating a new index and loading your data again. Document-schema indexes are also not covered by Pinecone's backup feature. And as of the day of writing, the command-line tool has no flag for creating a document index; use Python, TypeScript or the REST API.
For a first course, our recommendation is plain. Start with the integrated index to learn the ideas, learn the classic index because most code uses it, and keep the document index in mind for new projects that need keyword and vector search in one place.
pinecone.preview module that early document-index tutorials used. If you see pc.preview.indexes or from pinecone.preview import SchemaBuilder, drop the word preview and import from pinecone directly.- Create the
articlesindex with the schema above and upsert the two documents. - Run both searches and compare which scoring method fits which kind of question.
- Add
"year": 2021to a third document and write a filter that excludes it. - Delete the index.
Embeddings: where vectors come from
Until now the vectors were tiny numbers we invented, or Pinecone made them silently. In real projects you must understand this step, because the quality of your search is mostly the quality of your embeddings, not the database.
Pinecone hosts embedding models you can call directly. Calling one yourself, without an integrated index, looks like this:
from pinecone import Pinecone
pc = Pinecone()
out = pc.inference.embed(
model="llama-text-embed-v2",
inputs=["Reset my password", "Quarterly sales report"],
parameters={"input_type": "passage"},
)
print(len(out[0]["values"])) # 1024 by default
The hosted models you will see most often are these. llama-text-embed-v2 is a dense model that produces 1,024-number vectors by default (it also offers 2,048, 768, 512 and 384) and reads up to 2,048 tokens per input. multilingual-e5-large produces 1,024 dimensions and reads up to 507 tokens, and is a natural choice when your content mixes languages, for example Arabic and English. pinecone-sparse-english-v0 produces sparse vectors for keyword-style matching. A token is a chunk of text, roughly a short word or word fragment; models have a maximum number of tokens per input, and anything longer must be split first. Batches are limited to 96 inputs per call.
The input_type parameter is required and must be passage or query. These models are trained so that the text you store and the text people ask are embedded slightly differently, and mixing them up quietly lowers the quality of results. Use passage for the content you store and query for the user's question. An integrated index does this for you.
Whichever model you use, three rules prevent most embedding mistakes.
- Use the same model for documents and queries. Vectors from different models live in different spaces; comparing them is meaningless, and if the lengths differ you get a dimension error.
- Set the index dimension to the model's output length. Check the model card, not your memory.
- Use the metric the model was trained for. For most text models that is cosine.
Chunking is the step people skip and regret. A model has a token limit, and a single vector for a 30-page document blurs everything into one average meaning. So you split documents into passages, often a few hundred words each, and store each passage as its own record. Give related chunks IDs that share a prefix, such as handbook#chunk_0, handbook#chunk_1; the prefix lets you list and delete a whole document later. This is a design choice with no single right answer, and it is the main lever you will tune.
When your content is Arabic, check that the model you choose is documented to support it; a model trained mostly on English text will return poor results for Arabic even though the database works perfectly. The multilingual model above exists for exactly this reason.
Frameworks hide much of this plumbing. LangChain and LlamaIndex both have Pinecone integrations that embed, chunk and upsert for you. Learn the raw calls first so that when the framework misbehaves, you know which layer to inspect.
- Call
pc.inference.embedon two similar and one unrelated sentence and print the length of each vector. - Write down the chunk size you would choose for a 40-page PDF, and the ID scheme for its chunks.
- Say aloud why a query embedded with a different model would give meaningless scores.
Metadata and filtering
Similarity finds things that are alike. Filters find things that satisfy a rule. Real applications need both: "the five most relevant policy paragraphs, but only for the Egypt office and only from this year".
You attach metadata when you upsert, and you pass a filter when you search. Metadata is a flat JSON object. Allowed value types are strings, numbers, booleans and lists of strings. Nested objects and null values are not supported, and a whole record's filterable metadata may not exceed 40 KB. Flatten structures yourself: instead of {"office": {"country": "EG"}} store {"office_country": "EG"}.
A filter is a small JSON expression using operators that start with a dollar sign:
| Operator | Meaning | Example |
|---|---|---|
$eq, $ne |
equal, not equal | {"topic": {"$eq": "billing"}} |
$gt, $gte, $lt, $lte |
numeric comparison | {"year": {"$gte": 2024}} |
$in, $nin |
value in, or not in, a list | {"topic": {"$in": ["billing", "account"]}} |
$exists |
the field is present | {"year": {"$exists": true}} |
$and, $or, $not |
combine conditions | {"$and": [ ... ]} |
On a classic index the filter goes straight into query:
res = index.query(
namespace="demo",
vector=query_vector,
top_k=5,
filter={"$and": [{"topic": {"$eq": "billing"}}, {"year": {"$gte": 2024}}]},
include_metadata=True,
)
On an integrated index the filter goes inside the query object of search, as "filter": {...} next to top_k.
Filters are applied as part of the search, not after it. That matters: you ask for the top five and get five that satisfy the rule, rather than the top five overall with some removed. If fewer than five records pass the filter, you simply get fewer.
Some practical limits and habits:
- A single
$inor$nincan list at most 10,000 values. Longer lists return a400. - Writing
filter={}(an empty filter) is rejected by the new Python SDK before the request is even sent. Omit the argument when there is nothing to filter. - Do not put huge or free-text values in metadata. Keep metadata to the short facts you will filter on or display, and keep your full documents in your own database or object storage, storing only an ID that points back to them.
Good metadata
{"topic": "billing", "year": 2024, "lang": "ar"}- Short, flat, and useful for filtering
- An ID that points to the full text elsewhere
Bad metadata
{"doc": {"meta": {"year": 2024}}}(nested, rejected){"year": null}(null, rejected)- The whole 20-page document pasted in
- Using your six-sentence index from earlier, search "money" with a filter that keeps only
topic = billing. - Change the filter to
$inover two topics and see how the result set changes. - Write a filter that would find billing records from 2024 or later, using
$and.
Everyday operations: fetch, update, delete, list and count
Search is the headline feature, but day-to-day work is mostly maintenance: looking at a record, correcting it, removing it, and counting what you have. These calls are grouped by what you are trying to do. The examples use a classic index handle called index; the integrated and document styles have their own equivalents, which the official reference lists.
Look at specific records. fetch retrieves records by ID. It is exact, unlike search, and you can ask for up to 1,000 IDs at once:
got = index.fetch(ids=["a", "b"], namespace="shapes")
print(got.vectors["a"].metadata)
Change a record. upsert with an existing ID replaces the whole record. If you only want to alter a part, update changes the fields you name and leaves the rest:
index.update(id="a", set_metadata={"genre": "crimson"}, namespace="shapes")
You can also update many records at once by describing them with a metadata filter rather than listing IDs. That is a bulk operation with its own, lower rate limit, so reserve it for real bulk changes.
Remove records. delete takes IDs (up to 1,000 per call), a metadata filter, or a request to clear a whole namespace:
index.delete(ids=["c"], namespace="shapes")
index.delete(delete_all=True, namespace="shapes")
Deleting everything in a namespace is how you reset a tenant or a test. Deleting a namespace itself is also possible and is cheap. All deletes are permanent; there is no recycle bin.
delete_all=True works on the namespace you name, and the default namespace is a namespace like any other. Forgetting the namespace= argument means you operate on the default one. That is a common way to wipe the wrong data, or to find that a delete "did nothing" because your records live elsewhere.List IDs. If you use structured IDs such as handbook#chunk_0, list returns IDs that start with a prefix, up to 100 per page:
for page in index.list(prefix="handbook#", namespace="docs"):
print(page)
This is how you remove every chunk of one document: list by its prefix, then delete those IDs. It is the main reason to design IDs with a prefix on day one.
Count what you have. describe_index_stats reports the number of records per namespace and the total:
print(index.describe_index_stats())
Use it to confirm an upload worked: the count in the namespace should equal the number of records you sent. During the first seconds after a write the count may lag behind, for the same eventual-consistency reason as before.
Namespaces as a daily tool. To see which namespaces exist, call index.list_namespaces(). To create one before any data arrives, index.create_namespace(name="customer-42"). To remove one with all its records, index.delete_namespace(namespace="customer-42"). Namespace names can be up to 512 ASCII characters. Because a search only reads one namespace, a namespace per customer both isolates customers and keeps each customer's searches cheap and fast. Per-tenant indexes are the wrong tool: plans cap the number of indexes per project, while namespaces are plentiful.
Loading lots of data. A single upsert accepts at most 1,000 records or 2 MB, so large loads go in batches. The Python SDK can do the batching for you:
index.upsert(vectors=big_list, namespace="docs", batch_size=200)
On the current SDK, a bulk call keeps going when some batches fail and reports them rather than raising, so always inspect the result for failures after a large load: the response carries upserted_count, failed_item_count and errors. If you use the dataframe helper upsert_from_dataframe, install pandas yourself because it is not a dependency, and pass on_error="raise" if you prefer an exception to a quiet partial failure.
{source}#{chunk_number}. Then fetching, listing and deleting by document all become one-liners, and your metadata can stay small.- Upsert five records with IDs
guide#0toguide#4and three with prefixfaq#. - List by prefix
guide#, then delete exactly those IDs. - Run
describe_index_statsand check that only thefaq#records remain.
Using Pinecone from the command line, JavaScript and REST
Python is the most common choice, but Pinecone is an HTTP service and the other doors are worth knowing.
The CLI is good for quick inspection and for scripts. After pc auth login, you set a target organization and project, then work with indexes:
pc index create -n my-index -d 1536 -m cosine -c aws -r us-east-1
pc index list
pc index describe -i my-index
pc index stats -i my-index
pc index delete -i my-index
Notice the flags. -n names a new resource, -i points at an existing index, -d is the dimension, -m the metric, -c the cloud and -r the region. Add -j to get JSON output that another program can read. Since version 1.0, delete commands ask for confirmation, which will make a script hang in CI; pass --skip-confirmation (or --json) in automation. Old tutorials use top-level commands such as pc backup or pc collection; these now live under pc index backup and pc index collection.
JavaScript and TypeScript have an official package, @pinecone-database/pinecone. Version 9 requires Node.js 22 or newer, and the old docs page that says Node 18 is out of date. You install it with npm install @pinecone-database/pinecone, build the client with new Pinecone() (it reads the same PINECONE_API_KEY), and do server-side work only: never call Pinecone from browser code, which would expose your key and also fail with a cross-origin (CORS) error. REST is what every SDK calls underneath, which is useful for debugging. Two headers matter: Api-Key carries your key and X-Pinecone-Api-Version selects the API version. Always send the version header. If you omit it, the request quietly uses the oldest supported version, and when that version is retired your calls silently move to the next one, sometimes with breaking changes.
curl "https://api.pinecone.io/indexes" \
-H "Api-Key: $PINECONE_API_KEY" \
-H "X-Pinecone-Api-Version: 2026-07"
This is the control plane: the global endpoint at api.pinecone.io that creates, lists and deletes indexes. Reading and writing data goes through the data plane, which is a different address: the host of your particular index, for instance articles-abc123.svc.us-east-1.pinecone.io. You can read the host from pc index describe or from the console. This split explains a common confusion: your index "exists" according to the control plane, but your data call fails because you sent it to the wrong host.
2026-07, creating an index over raw REST takes a schema rather than the old top-level dimension and metric. If you want the classic request body, send the header X-Pinecone-Api-Version: 2026-04. The SDKs handle this for you, which is another reason to prefer them while learning.- Install the CLI, sign in, and run
pc index list. - Call the REST endpoint above with
curland compare its JSON with what Python printed. - Find your index's host in the output of
pc index describe.
Plans, limits and what things cost
Pinecone bills for what you store and what you do, so a beginner should understand the shape of the bill even on the free plan. These figures come from the official documentation at the time of writing and change; check the pricing page before relying on them.
The plans are Starter (free), Builder (a flat monthly fee, with operations blocked at quota rather than billed beyond it), Standard and Enterprise (monthly minimums, pay-as-you-go beyond that). The plans differ in how many projects, indexes, namespaces and users you can have, and in monthly allowances. On Starter you get one project, a handful of serverless indexes, a small storage allowance, and, importantly, only the us-east-1 region on AWS. Other regions need a paid plan. If a data-residency rule requires your data to stay in a particular geography, the region list matters: check which regions the docs currently offer before you design around them, because the Gulf and Egypt are not necessarily covered. For regulated data, ask your legal team before you create the index, since the region cannot be changed afterwards.
Billing uses three units. A read unit (RU) is charged for each query, scaled to the size of the namespace being searched. A write unit (WU) is charged for upserts, updates and deletes, scaled to the size of the data written. Storage is charged per gigabyte per month. There is also a charge for egress, the data your reads send back to you, above a monthly allowance.
Some consequences are worth absorbing now:
- A search costs more in a bigger namespace, because Pinecone reads more data. Splitting data into per-tenant namespaces lowers the cost of each search.
- Asking for
top_k=10ortop_k=1000costs the same number of read units, but returning more data, especially withinclude_values=Trueor heavy metadata, increases egress. - Idle indexes cost almost nothing beyond storage, but they still count towards your index limit. Delete the ones you do not need.
Pinecone also enforces rate limits, which matter in loops. Each namespace accepts about 100 requests per second for queries, upserts, updates and deletes. Above that you get a 429 response, covered in the error section below. And there are hard size limits that are easy to hit by accident: a record ID of at most 512 characters, filterable metadata of at most 40 KB, a dense dimension of at most 20,000, and a top_k of at most 10,000.
- Open the usage page in the console and note the read and write units after running your earlier scripts.
- Estimate the storage for one million records of 1,536 dimensions: each number takes 4 bytes, so multiply, then add your metadata size.
- List which of your own data would have to stay in a particular country, and check whether a nearby Pinecone region exists for it.
Configuration and security basics
There is not much to configure in Pinecone itself, which is part of its appeal, but a few settings decide whether your first project is safe and stable.
Credentials. Everything starts with the API key. Keys belong to a project and are shown once. Give each application its own key and delete keys you no longer use; deleting is immediate and irreversible, so anything still using that key stops working at once. On the paid plans you can restrict what a key may do (for example, read-only access to data), which is the right default for a service that only searches. Never commit keys, never log them, and never send them to a browser.
The Python client can take the key as an argument (Pinecone(api_key="...")), but reading it from the environment is better because your code then contains no secret. Other environment variables exist for advanced setups, and explicit arguments always win over them.
Version pinning. Pinecone's SDK versions are tied to API versions: Python 10 targets 2026-07, and upgrading the package is how you move to a new API version. In a real project pin the major version in your requirements file, for example pinecone>=10,<11, so that a surprise upgrade cannot change behaviour overnight. Read the migration notes when you do upgrade; the jump from version 8 to 9 and from 9 to 10 each removed or renamed several things.
Deletion protection. When you create an important index, turn it on:
pc.indexes.configure("articles", deletion_protection="enabled")
A delete on a protected index is refused with a 403 until you disable it. It is a seatbelt against the script that cleans up the wrong index.
- Create a second API key, use it for one script, then delete it in the console and watch that script fail with an authentication error.
- Add
pinecone>=10,<11to arequirements.txt. - Turn deletion protection on for a demo index, try to delete it, and read the refusal.
Reading errors: the common ones and what they mean
Errors from Pinecone are informative if you read the whole message. Below are the ones beginners actually meet, grouped by symptom.
Old code on a new library. Tutorials from 2023 start with pinecone.init(api_key=..., environment=...). That function was removed long ago, and you see AttributeError: module 'pinecone' has no attribute 'init'. The fix: from pinecone import Pinecone, then pc = Pinecone(api_key=...). There is no environment any more. Similarly, ImportError: cannot import name 'Pinecone' from 'pinecone' means an ancient installed version, or the old pinecone-client package. Run pip uninstall pinecone-client, then pip install --upgrade pinecone. If a message mentions a host like controller.us-west-2.pinecone.io, your library predates serverless; upgrade it.
Keyword mistakes from older versions. TypeError: Pinecone() got unexpected keyword arguments: ['retries'] means you used a retry keyword from version 8; the current form is retry_config=RetryConfig(...). TypeError: Pinecone.create_index() missing 1 required positional argument: 'spec' comes from versions that need a spec; pass spec=ServerlessSpec(...), or on version 10 use a schema and deployment.
HTTP status codes. The API replies with standard codes, and the SDK turns them into exceptions.
| Status | Meaning | What to do |
|---|---|---|
| 400 | The request is malformed: a dimension mismatch, too many $in values, a field type that violates the schema. |
Read the message; fix the request. |
| 401 | The key is missing, wrong, or belongs to another project. | Check the key and its project. |
| 403 | Forbidden: deletion protection, or a plan quota such as the maximum number of projects or indexes. | Disable protection deliberately, delete unused objects, or upgrade. |
| 404 | The index or resource does not exist under that name or host. | Check spelling, project and host. |
| 409 | You tried to create something that already exists. | Check with pc.has_index(name) first. |
| 412 | The operation is not valid in the current state, for example the index is not ready. | Wait, then retry. |
| 429 | Too many requests, or you hit a monthly quota. | Slow down, back off, spread load over namespaces, or upgrade. |
| 5xx | A problem on Pinecone's side. | Retry with backoff; check the status page. |
A 429 has several causes, and the message tells you which. "You've reached the query QPS limit for namespace..." is the per-second rate limit; pace your queries. "You've reached your read unit limit for the current month" is a plan quota that waits for the month to turn over or an upgrade.
Data that "is not there". The most frequent scare is empty search results right after a write. As covered, writes become searchable after a short delay. Run describe_index_stats to see whether the record count has caught up. A second cause is the wrong namespace: you wrote to demo and searched the default. A third is a metadata filter that is too strict. Remove the filter and see whether results appear.
Connection errors. Handshake read failed indicates a network, firewall, proxy or TLS problem. Test from another network, set ssl_ca_certs if your company intercepts TLS, and never disable verification. A browser console message about a missing Access-Control-Allow-Origin header means you called Pinecone from browser JavaScript; move the call to a backend.
- On purpose, upsert a vector of the wrong length into your three-dimension index and read the error.
- Try to create the same index name twice without
has_indexand note the status code you get. - Search a namespace that does not exist and observe what comes back.
Putting it all together: a small support-answer finder
Now combine everything into one small project: a helper that loads a handful of help-centre passages and answers questions by returning the best passage, restricted to a topic when the user asks. It is the retrieval half of a RAG system. Pairing it with a language model is the natural next step, and a web wrapper built with FastAPI would make it a service.
Plan before you code. The data is passages with a topic. We want a namespace per language so Arabic and English content never compete, an integrated index so we do not manage embeddings, meaningful IDs, and a cleanup step. Write support_finder.py:
import sys
import time
from pinecone import Pinecone
INDEX = "support-finder"
pc = Pinecone()
def ensure_index():
if not pc.has_index(INDEX):
pc.create_index_for_model(
name=INDEX, cloud="aws", region="us-east-1",
embed={"model": "multilingual-e5-large",
"field_map": {"text": "chunk_text"}},
)
return pc.index(INDEX)
def load(index):
en = [
{"_id": "en#0", "chunk_text": "Reset your password from Settings, then Security.", "topic": "account"},
{"_id": "en#1", "chunk_text": "Refunds reach your card within five days.", "topic": "billing"},
{"_id": "en#2", "chunk_text": "Invoices are emailed monthly.", "topic": "billing"},
]
index.upsert_records("en", en)
def ask(index, question, topic=None):
query = {"inputs": {"text": question}, "top_k": 2}
if topic:
query["filter"] = {"topic": {"$eq": topic}}
res = index.search(namespace="en", query=query,
fields=["chunk_text", "topic"])
return res["result"]["hits"]
if __name__ == "__main__":
idx = ensure_index()
if "--load" in sys.argv:
load(idx)
time.sleep(10)
for hit in ask(idx, "when do I get my money back?", topic="billing"):
print(round(hit["_score"], 3), hit["fields"]["chunk_text"])
if "--cleanup" in sys.argv:
pc.delete_index(INDEX)
Run it in stages: python support_finder.py --load first, then python support_finder.py for a second query without reloading, then python support_finder.py --cleanup to delete the index. Here the model is multilingual-e5-large, so the same index could hold an ar namespace of Arabic passages searched exactly the same way, with no changes except the namespace name. Notice how each section's ideas appear: a namespace keeps a language apart, the ID prefix names the language and chunk, metadata powers the filter, and the delay after writing is handled explicitly.
- Run the script in the three stages described and read each output.
- Add three Arabic passages to a namespace called
arand search them with an Arabic question. - Add a second filter on a new metadata field of your choice and test it.
What you can now do, and what comes next
You can now explain what Pinecone is for, name its five nouns, and describe the difference between a classic vector index, an integrated-embedding index and a document index. You can create an index, load records, wait for them to become searchable, search by meaning, filter on metadata, fetch, update, delete and count records, and read the common errors. You know that the cloud and region are permanent, that the dimension must match the model, that writes are eventually consistent, that API keys are secrets scoped to a project, and that you should send the API version header and pin the SDK.
Equally important, you know where this guide stops. Production questions are the next level: how to check that a write is visible, how to combine keyword and meaning search, how to use rerankers, how to batch and parallelise large loads, how to monitor usage, and how to design namespaces for many tenants. Beyond that sit the platform questions of backups, access control, private networking, cost control and when Pinecone is the wrong choice. Those appear in the Mid-level and Senior parts of this guide.
Meanwhile, a sensible practice path: build the support finder over your own documents; try the same data in Qdrant or Chroma to see how another vector database differs; pair retrieval with a framework such as LangChain; and later learn to measure retrieval quality with RAGAS, because a search that returns something is not the same as a search that returns the right thing.
Sources
- Pinecone documentation home
- Quickstart
- Key terms
- Create an index
- Adopt the Documents API
- Upsert data
- Semantic search
- Filter by metadata
- Manage namespaces
- Fetch, update and delete data
- Understanding cost
- API versioning
- Errors
- Database limits
- Models overview
- Python SDK installation
- Python SDK v10 migration
- CLI quickstart
- Local development with Pinecone Local