ElevenLabs Integration Guide
Technology: eleven-labs · Category: ai · Last reviewed: 2026-08-23
Source: https://tech-stack.codeamanilabs.org/guide/eleven-labs
Insight:
ElevenLabs adds voice — TTS for IVR / voice-note replies and STT for transcribing user audio. Pairs with Africa's Talking Voice and WhatsApp voice notes; multilingual voices matter for Swahili/Sheng audiences. Keep the API key server-side and stream audio to stay responsive on low bandwidth.
███████╗██╗ ███████╗██╗ ██╗███████╗███╗ ██╗██╗ █████╗ ██████╗ ███████╗
██╔════╝██║ ██╔════╝██║ ██║██╔════╝████╗ ██║██║ ██╔══██╗██╔══██╗██╔════╝
█████╗ ██║ █████╗ ██║ ██║█████╗ ██╔██╗ ██║██║ ███████║██████╔╝███████╗
██╔══╝ ██║ ██╔══╝ ╚██╗ ██╔╝██╔══╝ ██║╚██╗██║██║ ██╔══██║██╔══██╗╚════██║
███████╗███████╗███████╗ ╚████╔╝ ███████╗██║ ╚████║███████╗██║ ██║██████╔╝███████║
╚══════╝╚══════╝╚══════╝ ╚═══╝ ╚══════╝╚═╝ ╚═══╝╚══════╝╚═╝ ╚═╝╚═════╝ ╚══════╝
ElevenLabs Integration Guide
Focus: Text-to-speech, voice cloning, speech-to-text, and voice design via ElevenLabs APIs and Claude Code tooling.
Overview
ElevenLabs provides state-of-the-art AI voice generation — synthesize speech, clone voices, transcribe audio, and design custom voices. Used in codeAmani products for voice UI features, audio notifications, and multilingual TTS (including Swahili).
Here is the big picture — text flows out as audio, and user audio flows back in as text, all through one API:
flowchart LR
A["App text<br/>e.g. dashboard reply"] -->|"text_to_speech"| B["ElevenLabs API"]
B --> C["Audio stream<br/>mp3"]
C --> D["Play to user<br/>IVR or voice note"]
E["User audio<br/>recording"] -->|"speech_to_text<br/>scribe_v2"| B
B --> F["Transcribed text"]
Official Documentation
| Resource | URL |
|---|---|
| API Reference | https://elevenlabs.io/docs/api-reference |
| Node.js SDK | https://github.com/elevenlabs/elevenlabs-js |
| Python SDK | https://github.com/elevenlabs/elevenlabs-python |
| Voice Library | https://elevenlabs.io/voice-library |
| Models Reference | https://elevenlabs.io/docs/models |
MCP Server Setup
ElevenLabs offers two MCP paths. Prefer the hosted server — there is nothing to install and it authenticates over OAuth (no API key copied into a config file).
Note: there is no
@elevenlabs/elevenlabs-mcpnpm package (a common mistake —npxwill 404). The self-hosted server is the Python packageelevenlabs-mcp(PyPI), and its GitHub repo was archived on 2026-08-20 in favor of the hosted server below.
Hosted MCP (recommended — OAuth, no install)
# Streamable-HTTP remote server; completes an OAuth sign-in on first connect
claude mcp add --transport http elevenlabs https://api.elevenlabs.io/v1/mcp
The hosted server focuses on ElevenLabs Agents management (list/create/update agents in your workspace). Revoke access any time from your ElevenLabs account settings.
Self-hosted (local Python server — creative tools)
For the creative toolset (TTS, STT, voice management, sound generation) run the
elevenlabs-mcp PyPI package locally via uvx (requires the uv Python tool):
// .mcp.json
{
"mcpServers": {
"elevenlabs": {
"command": "uvx",
"args": ["elevenlabs-mcp"],
"env": {
"ELEVENLABS_API_KEY": "${ELEVENLABS_API_KEY}"
}
}
}
}
MCP Capabilities (local elevenlabs-mcp server)
| Tool | Description |
|---|---|
text_to_speech |
Convert text to audio using a specified voice |
list_voices |
Browse available and cloned voices |
get_voice |
Get details and settings for a voice |
speech_to_text |
Transcribe audio files |
sound_generation |
Generate sound effects from text |
SDK Setup
Node.js / TypeScript
The npm package was renamed — the old bare elevenlabs package is deprecated
("moved to @elevenlabs/elevenlabs-js"). Install the scoped package (current
@elevenlabs/elevenlabs-js is v2.64.0, SDK v2):
npm install @elevenlabs/elevenlabs-js
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
const client = new ElevenLabsClient({
apiKey: process.env.ELEVENLABS_API_KEY,
});
Python
The Python package keeps the bare elevenlabs name (current v2.64.0):
pip install elevenlabs
from elevenlabs.client import ElevenLabs
client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])
Core Patterns
You are about to wire these up — here is how a streaming TTS request travels from your Next.js route to the listener:
sequenceDiagram
participant Client
participant Route as "API route<br/>/api/tts"
participant EL as "ElevenLabs"
Client->>Route: POST text + voiceId
Route->>EL: textToSpeech.stream with modelId
EL-->>Route: audio chunks
Route-->>Client: audio/mpeg response
Client->>Client: play audio
Text-to-Speech (Streaming)
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });
// Stream audio to a file (v2 SDK: textToSpeech.stream, camelCase params)
const audioStream = await client.textToSpeech.stream(
"JBFqnCBsd6RMkjVDRZzb", // Voice ID (Rachel — default)
{
text: "Karibu! Welcome to your codeAmani dashboard.",
modelId: "eleven_multilingual_v2",
voiceSettings: {
stability: 0.5,
similarityBoost: 0.75,
},
}
);
// Write to disk
import { createWriteStream } from "fs";
const writer = createWriteStream("output.mp3");
for await (const chunk of audioStream) {
writer.write(chunk);
}
writer.end();
Next.js App Router API Route
// app/api/tts/route.ts
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { NextRequest } from "next/server";
const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });
export async function POST(req: NextRequest) {
const { text, voiceId = "JBFqnCBsd6RMkjVDRZzb" } = await req.json();
const audioStream = await client.textToSpeech.stream(voiceId, {
text,
modelId: "eleven_multilingual_v2",
});
// Collect chunks
const chunks: Buffer[] = [];
for await (const chunk of audioStream) {
chunks.push(Buffer.from(chunk));
}
return new Response(Buffer.concat(chunks), {
headers: {
"Content-Type": "audio/mpeg",
"Cache-Control": "no-store",
},
});
}
List Available Voices
// v2 SDK: voices.search() (GET /v2/voices, paginated). voices.getAll() still
// works as a legacy alias but search() is the current method.
const { voices } = await client.voices.search();
for (const voice of voices) {
console.log(`${voice.voice_id}: ${voice.name} (${voice.labels?.language ?? "multi"})`);
}
Speech-to-Text (Transcription)
import { createReadStream } from "fs";
const transcription = await client.speechToText.convert({
file: createReadStream("recording.mp3"),
modelId: "scribe_v2", // current STT model (scribe_v1 is deprecated)
languageCode: "sw", // Swahili
});
console.log(transcription.text);
Voice cloning
Instant Voice Cloning (IVC) turns a short audio sample into a reusable voice. You upload one or more recordings, get back a voice_id, then synthesize speech with it like any other voice. This is how you give a codeAmani product its own branded voice — or let a user respond in their own voice for voice-note replies.
The flow is two API calls — clone once, reuse the voice_id forever:
flowchart LR
A["Audio samples<br/>clean · 1+ min"] -->|"voices.ivc.create"| B["ElevenLabs"]
B --> C["voice_id<br/>saved to account"]
C -->|"textToSpeech.convert"| D["Synthesized audio<br/>in cloned voice"]
C --> E["Store voice_id<br/>in your DB"]
Create a cloned voice, then synthesize with it
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import fs from "fs";
const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });
// 1. Clone a voice from one or more audio samples
const cloned = await client.voices.ivc.create({
name: "codeAmani Brand Voice",
files: [
fs.createReadStream("samples/sample-1.mp3"),
fs.createReadStream("samples/sample-2.mp3"),
],
});
const voiceId = cloned.voice_id;
console.log("Cloned voice id:", voiceId);
// Persist voiceId in your DB so you can reuse it without re-cloning.
// 2. Synthesize speech using the returned voice_id
const audio = await client.textToSpeech.convert(voiceId, {
text: "Karibu! Hii ni sauti yako mpya kwenye codeAmani.",
modelId: "eleven_multilingual_v2",
outputFormat: "mp3_44100_128",
});
const chunks: Buffer[] = [];
for await (const chunk of audio) {
chunks.push(Buffer.from(chunk));
}
fs.writeFileSync("cloned-output.mp3", Buffer.concat(chunks));
The create call returns an AddVoiceIVCResponseModel with voice_id (use this everywhere) and requires_verification (whether the voice must be verified before high-volume use). The cloned voice also appears in client.voices.getAll().
Gotcha — consent and sample quality. Only clone voices you have explicit permission to use; cloning a real person's voice without consent breaks ElevenLabs' terms (and KDPA-style consent expectations for biometric/voice data). Quality is bounded by your samples: use clean, single-speaker recordings with no background noise or music — a minute of crisp audio beats ten minutes of noisy phone audio. Keep the API key server-side; never expose it to the client doing the upload.
Model Reference
| Model ID | Use Case | Languages | Latency |
|---|---|---|---|
eleven_v3 |
Flagship — most expressive/emotional TTS, long-form | 70+ | higher |
eleven_v3_conversational |
Expressive real-time TTS for voice agents | 70+ | ~280ms |
eleven_multilingual_v2 |
Lifelike, consistent high-quality TTS | 29 | ~1s |
eleven_flash_v2_5 |
Lowest latency / cheapest (50% off per char), bulk & real-time | 32 | ~75ms |
scribe_v2 |
Speech-to-text transcription (batch) | 90+ | — |
scribe_v2_realtime |
Streaming speech-to-text | 90+ | ~150ms |
Deprecated / superseded (do not use in new code):
eleven_turbo_v2_5andeleven_turbo_v2— ElevenLabs recommends the Flash models in all use cases (functionally equivalent, lower latency).scribe_v1→ migrate toscribe_v2.For Swahili TTS,
eleven_multilingual_v2remains the proven choice for quality;eleven_v3(70+ languages) is the newer, more expressive option andeleven_flash_v2_5(32 languages) covers low-latency/real-time Swahili.
Webhooks
ElevenLabs sends webhook events for async operations (e.g. batch speech-to-text,
post-call transcripts). The signature is HMAC-SHA256 in the ElevenLabs-Signature
header, formatted t=<unix_ts>,v0=<hex_hmac>, where the signed message is
`${timestamp}.${rawBody}` (not the raw body alone). Reject requests whose
timestamp is outside a ~30-minute window to block replays.
Prefer the SDK helper, which verifies the signature, checks the timestamp window, and parses the payload for you:
// app/api/webhooks/elevenlabs/route.ts
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { NextRequest } from "next/server";
const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });
export async function POST(req: NextRequest) {
const body = await req.text();
const sigHeader = req.headers.get("ElevenLabs-Signature") ?? "";
const secret = process.env.ELEVENLABS_WEBHOOK_SECRET!;
let event;
try {
event = await client.webhooks.constructEvent(body, sigHeader, secret);
} catch {
return new Response("Unauthorized", { status: 401 });
}
// Handle event.type, e.g. "speech_to_text_transcription.completed"
return new Response("OK");
}
If you verify manually, recompute HMAC_SHA256(secret, ${t}.${body}) and compare
(constant-time) against the v0= value parsed from the header.
Environment Variables
# Required
ELEVENLABS_API_KEY=sk_...
# Optional
ELEVENLABS_WEBHOOK_SECRET=whsec_...
ELEVENLABS_DEFAULT_VOICE_ID=JBFqnCBsd6RMkjVDRZzb
Common Use Cases
| Use Case | Approach |
|---|---|
| Voice notifications | POST to /api/tts, play on client |
| Swahili audio content | eleven_multilingual_v2 + languageCode: "sw" |
| Voice cloning | voices.ivc.create() with audio samples → get voice_id → use in TTS (see Voice cloning) |
| Audio transcription | scribe_v2 model + speechToText.convert() |
| Real-time voice AI | WebSocket streaming with eleven_flash_v2_5 (or eleven_v3_conversational) |
Troubleshooting
| Issue | Fix |
|---|---|
401 Unauthorized |
Check ELEVENLABS_API_KEY is set and valid |
| Audio sounds robotic | Increase stability (0.7–0.9) and similarity_boost (0.8) |
| Swahili not accurate | Use eleven_multilingual_v2 (or eleven_v3); Flash trades some quality for latency |
| Large audio files slow | Use streaming (textToSpeech.stream) instead of buffered response |
| MCP server not connecting | Run claude mcp list and check ELEVENLABS_API_KEY in env |
Official docs: