Pinecone Integration Guide

Technology: pinecone · Category: database · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/pinecone

Insight:

This directly solves the stated gap ("lots of AI, no retrieval layer"). Pinecone's integrated inference is the key: createIndexForModel binds an embedding model to the index, so you upsertRecords/searchRecords with raw text — Pinecone embeds server-side, meaning zero separate embeddings infrastructure to stand up. Two codeAmani wins: Pinecone's hosted multilingual-e5-large embeds Swahili + English in the same space (East-African content works out of the box), and retrieved chunks feed straight into a Claude call via the Vercel AI SDK — a full RAG loop across two guides.

██████╗ ██╗███╗   ██╗███████╗ ██████╗ ██████╗ ███╗   ██╗███████╗
██╔══██╗██║████╗  ██║██╔════╝██╔════╝██╔═══██╗████╗  ██║██╔════╝
██████╔╝██║██╔██╗ ██║█████╗  ██║     ██║   ██║██╔██╗ ██║█████╗
██╔═══╝ ██║██║╚██╗██║██╔══╝  ██║     ██║   ██║██║╚██╗██║██╔══╝
██║     ██║██║ ╚████║███████╗╚██████╗╚██████╔╝██║ ╚████║███████╗
╚═╝     ╚═╝╚═╝  ╚═══╝╚══════╝ ╚═════╝ ╚═════╝ ╚═╝  ╚═══╝╚══════╝

Pinecone Integration Guide

Focus: The retrieval layer for codeAmani's AI. Pinecone is a serverless vector database for RAG — store knowledge, retrieve the most relevant chunks, and feed them into a Claude prompt. With integrated inference you skip running an embedding model entirely; Pinecone embeds text server-side (incl. multilingual-e5-large for Swahili+English).

Overview

Two ways to use Pinecone — pick based on whether you want Pinecone to do the embedding:

Namespaces partition an index (e.g. per tenant/user) — the multi-tenant isolation primitive. In Claude Code, the Pinecone MCP (npx -y @pinecone-database/mcp) can search the docs, create indexes, and upsert/search records directly — see §5.

Here is the integrated-inference path at a glance — notice the embedding happens server-side, so your app never touches a vector:

flowchart LR
  A["Raw text records"] -->|"upsertRecords"| B["Index bound to<br/>multilingual-e5-large"]
  B --> C["Vectors stored<br/>server-side"]
  Q["Text query"] -->|"searchRecords topK"| C
  C --> D["Optional rerank<br/>bge-reranker-v2-m3 topN"]
  D --> E["Top matching chunks"]

RAG decision notes

The headline trade-off — Pinecone's integrated inference vs. running your own embeddings — is summarized in the Insight callout at the top of this page. A few more decision points worth knowing:

Chunking strategy

Since chunking dominates retrieval quality (see the note above), here is concrete guidance — all sizes are in tokens, not characters.

Size — start fixed, then iterate. Pinecone's chunking guide recommends starting with fixed-size chunking and only moving to fancier strategies once it proves insufficient. For sizing, the guide says to "start by exploring a variety of chunk sizes, including smaller chunks (e.g. 128 or 256 tokens) ... and larger chunks (e.g. 512 or 1024 tokens)" (pinecone.io/learn/chunking-strategies). Practical default for FAQ/support docs: ~512 tokens. Smaller chunks give sharper, more precise matches; larger chunks carry more context per hit but dilute the embedding.

Overlap — 10–20% is the commonly-documented range. Pinecone's guide itself does not prescribe a fixed overlap percentage; it instead favours chunk expansion (pulling neighbouring chunks at query time) to recover context. Across the broader RAG literature, a 10–20% overlap (e.g. 50–100 tokens on a 512-token chunk) is the widely-cited starting point to avoid splitting a sentence's meaning across a boundary. Treat it as a knob to tune, not gospel — some recent benchmarks find overlap adds indexing cost with little gain, so measure on your own corpus.

Semantic vs fixed splitting — when to switch:

Keep chunk_text (the field in your index fieldMap) as the only text field per record, and store source metadata (doc_id, section, chunk_index) so you can rebuild context.

// Fixed-size, overlapping chunker — run BEFORE upsertRecords (§2).
// Approximate tokens with a chars-per-token ratio (English ~4; tune for Swahili).
const CHARS_PER_TOKEN = 4;

function chunkText(
  text: string,
  { chunkTokens = 512, overlapTokens = 80 } = {},
): string[] {
  const size = chunkTokens * CHARS_PER_TOKEN;        // ~2048 chars
  const overlap = overlapTokens * CHARS_PER_TOKEN;    // ~320 chars
  const stride = Math.max(size - overlap, 1);         // guard: overlap < size
  const chunks: string[] = [];
  for (let start = 0; start < text.length; start += stride) {
    const slice = text.slice(start, start + size).trim();
    if (slice) chunks.push(slice);
  }
  return chunks;
}

// Wire chunks into the §2 integrated-inference upsert:
const records = chunkText(sourceDoc).map((chunk, i) => ({
  id: `faq-mpesa#${i}`,           // stable id = doc + chunk index
  chunk_text: chunk,             // must match the index fieldMap
  doc_id: "faq-mpesa",
  chunk_index: i,
  topic: "payments",
}));
await index.upsertRecords({ records });

Gotcha: chunk sizes are measured in tokens, but most splitters (including the char-based helper above) cut on characters. The ~4-chars-per-token ratio is an English approximation — Swahili and other non-English text often run fewer chars per token, so a "512-token" char window can silently overshoot the embedding model's context limit and get truncated server-side. For production, count with a real tokenizer (e.g. tiktoken / js-tiktoken) instead of a fixed ratio, and verify against your model's window (multilingual-e5-large ≈ 507 tokens; llama-text-embed-v2 = 2048 tokens).

flowchart TD
  A["Source doc"] --> B{"Structured doc<br/>headings · sections"}
  B -->|"No · uniform prose"| C["Fixed-size<br/>~512 tok · 10-20% overlap"]
  B -->|"Yes"| D{"Topics shift<br/>a lot"}
  D -->|"No"| E["Content-aware<br/>recursive split"]
  D -->|"Yes · long PDF"| F["Semantic<br/>split on topic shift"]
  C --> G["upsertRecords<br/>chunk_text field"]
  E --> G
  F --> G

Official Documentation

Resource URL
Docs home https://docs.pinecone.io/
Quickstart https://docs.pinecone.io/guides/get-started/quickstart
Integrated inference https://docs.pinecone.io/guides/inference/understanding-inference
Chunking strategies https://www.pinecone.io/learn/chunking-strategies/
TypeScript client https://github.com/pinecone-io/pinecone-ts-client
Python SDK https://github.com/pinecone-io/python-sdk

1. Install + credentials

npm install @pinecone-database/pinecone     # JS/TS — v8.x
pip install pinecone                         # Python 3.10+ — v9.x
PINECONE_API_KEY=...    # server-side only (.env.local / Vercel env, or Infisical)

Optional — the pc CLI (for scripting index/project ops outside the app):

brew install pinecone-io/tap/pinecone       # macOS/Linux (Homebrew)
# or: curl -fsSL https://pinecone.io/install.sh | sh
pc auth login                                # browser auth
pc index create --name kb --dimension 1024 --metric cosine --cloud aws --region us-east-1
pc index list

2. RAG with integrated inference (no embedding layer) — TypeScript

import { Pinecone } from "@pinecone-database/pinecone";

const pc = new Pinecone({ apiKey: process.env.PINECONE_API_KEY! });

// 2a. Create an index bound to an embedding model (multilingual = Swahili + English)
const model = await pc.createIndexForModel({
  name: "kb",
  cloud: "aws",
  region: "us-east-1",
  embed: { model: "multilingual-e5-large", fieldMap: { text: "chunk_text" } },
  waitUntilReady: true,
});
const index = pc.index({ host: model.host });

// 2b. Upsert RAW TEXT — Pinecone embeds it server-side
await index.upsertRecords({
  records: [
    { id: "doc1", chunk_text: "M-Pesa STK Push prompts the user on their phone.", topic: "payments" },
    { id: "doc2", chunk_text: "Daraja tokens expire after one hour.", topic: "payments" },
  ],
});

// 2c. Search with a text query (+ optional rerank)
const hits = await index.searchRecords({
  query: { topK: 4, inputs: { text: "how does M-Pesa checkout work?" }, filter: { topic: { $eq: "payments" } } },
  rerank: { model: "bge-reranker-v2-m3", topN: 2, rankFields: ["chunk_text"] },
  fields: ["chunk_text", "topic"],
});

3. Close the RAG loop (Pinecone → Claude via the AI SDK)

You are one short hop from a full answer — retrieved chunks become Claude's context, with Pinecone as memory and Claude as generator:

sequenceDiagram
  participant U as "User"
  participant App as "Server action"
  participant P as "Pinecone"
  participant C as "Claude · AI SDK"
  U->>App: "Question"
  App->>P: "searchRecords · text query"
  P-->>App: "Top-k chunks"
  App->>C: "Prompt with context"
  C-->>App: "Grounded answer"
  App-->>U: "Reply"
import { generateText } from "ai";

const context = hits.result.hits.map(h => h.fields.chunk_text).join("\n---\n");
const { text } = await generateText({
  model: "anthropic/claude-sonnet-5",               // via Vercel AI Gateway (Claude 5 family)
  prompt: `Answer using ONLY this context:\n${context}\n\nQ: How does M-Pesa checkout work?`,
});

4. Bring-your-own-vectors (Python, explicit dimension)

from pinecone import Pinecone, ServerlessSpec
pc = Pinecone(api_key="...")
pc.indexes.create(name="kb", dimension=1536, metric="cosine",
                  spec=ServerlessSpec(cloud="aws", region="us-east-1"))
index = pc.index("kb")   # v9: lowercase pc.index(); capital pc.Index() is deprecated
index.upsert(vectors=[("id1", embedding_1536d)], namespace="tenant-a")
res = index.query(vector=query_vec, top_k=10, namespace="tenant-a")

5. Claude Code — Pinecone MCP + pc CLI

The Pinecone MCP (npx -y @pinecone-database/mcp, PINECONE_API_KEY in the env) lets an agent build and inspect indexes without leaving the editor. Current tools:

Tool Use
search-docs Look up current Pinecone docs
list-indexes / describe-index / describe-index-stats Inspect indexes + record counts
create-index-for-model Create an integrated-inference index
upsert-records / search-records Write and query raw-text records
rerank-documents Rerank a candidate list
cascading-search Search across multiple indexes and merge

The MCP supports integrated-embedding indexes only — bring-your-own-vector indexes are not addressable through it. For scripting and CI, the standalone pc CLI (see §1) covers auth, project targeting, and index/vector CRUD.

codeAmani notes

Official docs: