Pinecone Integration Guide
Technology: pinecone · Category: database · Last reviewed: 2026-08-23
Source: https://tech-stack.codeamanilabs.org/guide/pinecone
Insight:
This directly solves the stated gap ("lots of AI, no retrieval layer"). Pinecone's integrated inference is the key:
createIndexForModelbinds an embedding model to the index, so youupsertRecords/searchRecordswith raw text — Pinecone embeds server-side, meaning zero separate embeddings infrastructure to stand up. Two codeAmani wins: Pinecone's hostedmultilingual-e5-largeembeds Swahili + English in the same space (East-African content works out of the box), and retrieved chunks feed straight into a Claude call via the Vercel AI SDK — a full RAG loop across two guides.
██████╗ ██╗███╗ ██╗███████╗ ██████╗ ██████╗ ███╗ ██╗███████╗
██╔══██╗██║████╗ ██║██╔════╝██╔════╝██╔═══██╗████╗ ██║██╔════╝
██████╔╝██║██╔██╗ ██║█████╗ ██║ ██║ ██║██╔██╗ ██║█████╗
██╔═══╝ ██║██║╚██╗██║██╔══╝ ██║ ██║ ██║██║╚██╗██║██╔══╝
██║ ██║██║ ╚████║███████╗╚██████╗╚██████╔╝██║ ╚████║███████╗
╚═╝ ╚═╝╚═╝ ╚═══╝╚══════╝ ╚═════╝ ╚═════╝ ╚═╝ ╚═══╝╚══════╝
Pinecone Integration Guide
Focus: The retrieval layer for codeAmani's AI. Pinecone is a serverless vector database for RAG — store knowledge, retrieve the most relevant chunks, and feed them into a Claude prompt. With integrated inference you skip running an embedding model entirely; Pinecone embeds text server-side (incl.
multilingual-e5-largefor Swahili+English).
Overview
Two ways to use Pinecone — pick based on whether you want Pinecone to do the embedding:
- Integrated inference (recommended for RAG):
createIndexForModelbinds an embedding model to the index. You thenupsertRecords/searchRecordswith raw text — no separate embeddings step, no vector math in your app. Pinecone's current default hosted model isllama-text-embed-v2(1024-dim, 2048-token window); codeAmani deliberately picksmultilingual-e5-large(1024-dim, ~507-token window) for its Swahili + English coverage. Optional server-side reranking (bge-reranker-v2-m3) sharpens results. - Bring-your-own-vectors: create a serverless index with an explicit
dimension, embed text yourself (e.g. via the Vercel AI SDK / Gemini / OpenAI), thenupsert/queryraw vectors. Use when you need a specific embedding model.
Namespaces partition an index (e.g. per tenant/user) — the multi-tenant isolation
primitive. In Claude Code, the Pinecone MCP (npx -y @pinecone-database/mcp) can search
the docs, create indexes, and upsert/search records directly — see §5.
Here is the integrated-inference path at a glance — notice the embedding happens server-side, so your app never touches a vector:
flowchart LR
A["Raw text records"] -->|"upsertRecords"| B["Index bound to<br/>multilingual-e5-large"]
B --> C["Vectors stored<br/>server-side"]
Q["Text query"] -->|"searchRecords topK"| C
C --> D["Optional rerank<br/>bge-reranker-v2-m3 topN"]
D --> E["Top matching chunks"]
RAG decision notes
The headline trade-off — Pinecone's integrated inference vs. running your own embeddings — is summarized in the Insight callout at the top of this page. A few more decision points worth knowing:
- When NOT to use integrated inference: if you need a specific embedding model (e.g. to match vectors you already store elsewhere, or a domain-tuned model), use the bring-your-own-vectors path (§4) and embed with the AI SDK instead.
- Reranking earns its cost: a 2-stage retrieve→rerank (
bge-reranker-v2-m3) noticeably lifts answer quality for FAQ/support RAG — keeptopKwide (e.g. 20) andtopNtight (e.g. 3). - Chunking matters more than the DB: retrieval quality is dominated by how you split source docs (size + overlap), not by Pinecone config — tune chunks first.
- Serverless = scale-to-zero: no idle index cost, so spinning up a per-product knowledge base is cheap for early-stage SME features.
Chunking strategy
Since chunking dominates retrieval quality (see the note above), here is concrete guidance — all sizes are in tokens, not characters.
Size — start fixed, then iterate. Pinecone's chunking guide recommends starting with fixed-size chunking and only moving to fancier strategies once it proves insufficient. For sizing, the guide says to "start by exploring a variety of chunk sizes, including smaller chunks (e.g. 128 or 256 tokens) ... and larger chunks (e.g. 512 or 1024 tokens)" (pinecone.io/learn/chunking-strategies). Practical default for FAQ/support docs: ~512 tokens. Smaller chunks give sharper, more precise matches; larger chunks carry more context per hit but dilute the embedding.
Overlap — 10–20% is the commonly-documented range. Pinecone's guide itself does not prescribe a fixed overlap percentage; it instead favours chunk expansion (pulling neighbouring chunks at query time) to recover context. Across the broader RAG literature, a 10–20% overlap (e.g. 50–100 tokens on a 512-token chunk) is the widely-cited starting point to avoid splitting a sentence's meaning across a boundary. Treat it as a knob to tune, not gospel — some recent benchmarks find overlap adds indexing cost with little gain, so measure on your own corpus.
Semantic vs fixed splitting — when to switch:
- Fixed-size (token windows + overlap): deterministic, fast, zero document understanding. Use as your default and for uniform text (support tickets, chat logs, plain prose).
- Content-aware / recursive (split on
\n\n,\n, sentences, then fall back): respects structure. Use for Markdown/HTML docs, USSD scripts, code where paragraph and heading boundaries are meaningful. - Semantic (embed sentences, cut where the topic shifts): highest quality, highest cost. Use only for long, topic-switching documents (multi-section policy/compliance PDFs) where fixed splitting visibly hurts answer quality.
Keep chunk_text (the field in your index fieldMap) as the only text field per record,
and store source metadata (doc_id, section, chunk_index) so you can rebuild context.
// Fixed-size, overlapping chunker — run BEFORE upsertRecords (§2).
// Approximate tokens with a chars-per-token ratio (English ~4; tune for Swahili).
const CHARS_PER_TOKEN = 4;
function chunkText(
text: string,
{ chunkTokens = 512, overlapTokens = 80 } = {},
): string[] {
const size = chunkTokens * CHARS_PER_TOKEN; // ~2048 chars
const overlap = overlapTokens * CHARS_PER_TOKEN; // ~320 chars
const stride = Math.max(size - overlap, 1); // guard: overlap < size
const chunks: string[] = [];
for (let start = 0; start < text.length; start += stride) {
const slice = text.slice(start, start + size).trim();
if (slice) chunks.push(slice);
}
return chunks;
}
// Wire chunks into the §2 integrated-inference upsert:
const records = chunkText(sourceDoc).map((chunk, i) => ({
id: `faq-mpesa#${i}`, // stable id = doc + chunk index
chunk_text: chunk, // must match the index fieldMap
doc_id: "faq-mpesa",
chunk_index: i,
topic: "payments",
}));
await index.upsertRecords({ records });
Gotcha: chunk sizes are measured in tokens, but most splitters (including the char-based helper above) cut on characters. The ~4-chars-per-token ratio is an English approximation — Swahili and other non-English text often run fewer chars per token, so a "512-token" char window can silently overshoot the embedding model's context limit and get truncated server-side. For production, count with a real tokenizer (e.g.
tiktoken/js-tiktoken) instead of a fixed ratio, and verify against your model's window (multilingual-e5-large≈ 507 tokens;llama-text-embed-v2= 2048 tokens).
flowchart TD
A["Source doc"] --> B{"Structured doc<br/>headings · sections"}
B -->|"No · uniform prose"| C["Fixed-size<br/>~512 tok · 10-20% overlap"]
B -->|"Yes"| D{"Topics shift<br/>a lot"}
D -->|"No"| E["Content-aware<br/>recursive split"]
D -->|"Yes · long PDF"| F["Semantic<br/>split on topic shift"]
C --> G["upsertRecords<br/>chunk_text field"]
E --> G
F --> G
Official Documentation
| Resource | URL |
|---|---|
| Docs home | https://docs.pinecone.io/ |
| Quickstart | https://docs.pinecone.io/guides/get-started/quickstart |
| Integrated inference | https://docs.pinecone.io/guides/inference/understanding-inference |
| Chunking strategies | https://www.pinecone.io/learn/chunking-strategies/ |
| TypeScript client | https://github.com/pinecone-io/pinecone-ts-client |
| Python SDK | https://github.com/pinecone-io/python-sdk |
1. Install + credentials
npm install @pinecone-database/pinecone # JS/TS — v8.x
pip install pinecone # Python 3.10+ — v9.x
PINECONE_API_KEY=... # server-side only (.env.local / Vercel env, or Infisical)
Optional — the pc CLI (for scripting index/project ops outside the app):
brew install pinecone-io/tap/pinecone # macOS/Linux (Homebrew)
# or: curl -fsSL https://pinecone.io/install.sh | sh
pc auth login # browser auth
pc index create --name kb --dimension 1024 --metric cosine --cloud aws --region us-east-1
pc index list
2. RAG with integrated inference (no embedding layer) — TypeScript
import { Pinecone } from "@pinecone-database/pinecone";
const pc = new Pinecone({ apiKey: process.env.PINECONE_API_KEY! });
// 2a. Create an index bound to an embedding model (multilingual = Swahili + English)
const model = await pc.createIndexForModel({
name: "kb",
cloud: "aws",
region: "us-east-1",
embed: { model: "multilingual-e5-large", fieldMap: { text: "chunk_text" } },
waitUntilReady: true,
});
const index = pc.index({ host: model.host });
// 2b. Upsert RAW TEXT — Pinecone embeds it server-side
await index.upsertRecords({
records: [
{ id: "doc1", chunk_text: "M-Pesa STK Push prompts the user on their phone.", topic: "payments" },
{ id: "doc2", chunk_text: "Daraja tokens expire after one hour.", topic: "payments" },
],
});
// 2c. Search with a text query (+ optional rerank)
const hits = await index.searchRecords({
query: { topK: 4, inputs: { text: "how does M-Pesa checkout work?" }, filter: { topic: { $eq: "payments" } } },
rerank: { model: "bge-reranker-v2-m3", topN: 2, rankFields: ["chunk_text"] },
fields: ["chunk_text", "topic"],
});
3. Close the RAG loop (Pinecone → Claude via the AI SDK)
You are one short hop from a full answer — retrieved chunks become Claude's context, with Pinecone as memory and Claude as generator:
sequenceDiagram
participant U as "User"
participant App as "Server action"
participant P as "Pinecone"
participant C as "Claude · AI SDK"
U->>App: "Question"
App->>P: "searchRecords · text query"
P-->>App: "Top-k chunks"
App->>C: "Prompt with context"
C-->>App: "Grounded answer"
App-->>U: "Reply"
import { generateText } from "ai";
const context = hits.result.hits.map(h => h.fields.chunk_text).join("\n---\n");
const { text } = await generateText({
model: "anthropic/claude-sonnet-5", // via Vercel AI Gateway (Claude 5 family)
prompt: `Answer using ONLY this context:\n${context}\n\nQ: How does M-Pesa checkout work?`,
});
4. Bring-your-own-vectors (Python, explicit dimension)
from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key="...")
pc.indexes.create(name="kb", dimension=1536, metric="cosine",
spec=ServerlessSpec(cloud="aws", region="us-east-1"))
index = pc.index("kb") # v9: lowercase pc.index(); capital pc.Index() is deprecated
index.upsert(vectors=[("id1", embedding_1536d)], namespace="tenant-a")
res = index.query(vector=query_vec, top_k=10, namespace="tenant-a")
5. Claude Code — Pinecone MCP + pc CLI
The Pinecone MCP (npx -y @pinecone-database/mcp, PINECONE_API_KEY in the env) lets an
agent build and inspect indexes without leaving the editor. Current tools:
| Tool | Use |
|---|---|
search-docs |
Look up current Pinecone docs |
list-indexes / describe-index / describe-index-stats |
Inspect indexes + record counts |
create-index-for-model |
Create an integrated-inference index |
upsert-records / search-records |
Write and query raw-text records |
rerank-documents |
Rerank a candidate list |
cascading-search |
Search across multiple indexes and merge |
The MCP supports integrated-embedding indexes only — bring-your-own-vector indexes are not
addressable through it. For scripting and CI, the standalone pc CLI (see §1) covers auth,
project targeting, and index/vector CRUD.
codeAmani notes
- Fills the retrieval gap: this is the missing layer for RAG over codeAmani docs/FAQs/
product knowledge. Retrieve top-k chunks, then pass them to Claude via the Vercel AI SDK
(see
vercel-ai-sdk). Keep Claude as the generator; Pinecone as the memory. - Multilingual by choice: codeAmani picks
multilingual-e5-large(Pinecone's default is nowllama-text-embed-v2) because it embeds Swahili and English in one space — East-African content (mixed-language support FAQs, USSD scripts) works without a custom model. Strong fit for theAFRICAN_MARKET_GUIDE.mdaudience. - Multi-tenant: use a namespace per customer/SME for isolation (RLS-like separation).
- Security:
PINECONE_API_KEYis server-side only (.env.local/ Vercel env, or move it into Infisical). Query from API routes / server actions, never the browser. - Cost: serverless pay-per-use; reads cost readUnits and reranking costs rerankUnits —
cache hot queries (pairs with the Upstash cache) and keep
topK/topNtight. - Claude Code: the Pinecone MCP (
@pinecone-database/mcp) and thepcCLI let you build and inspect indexes interactively without leaving the editor — see §5.
Official docs: