ElevenLabs Integration Guide

Technology: eleven-labs · Category: ai · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/eleven-labs

Insight:

ElevenLabs adds voice — TTS for IVR / voice-note replies and STT for transcribing user audio. Pairs with Africa's Talking Voice and WhatsApp voice notes; multilingual voices matter for Swahili/Sheng audiences. Keep the API key server-side and stream audio to stay responsive on low bandwidth.

███████╗██╗     ███████╗██╗   ██╗███████╗███╗   ██╗██╗      █████╗ ██████╗ ███████╗
██╔════╝██║     ██╔════╝██║   ██║██╔════╝████╗  ██║██║     ██╔══██╗██╔══██╗██╔════╝
█████╗  ██║     █████╗  ██║   ██║█████╗  ██╔██╗ ██║██║     ███████║██████╔╝███████╗
██╔══╝  ██║     ██╔══╝  ╚██╗ ██╔╝██╔══╝  ██║╚██╗██║██║     ██╔══██║██╔══██╗╚════██║
███████╗███████╗███████╗ ╚████╔╝ ███████╗██║ ╚████║███████╗██║  ██║██████╔╝███████║
╚══════╝╚══════╝╚══════╝  ╚═══╝  ╚══════╝╚═╝  ╚═══╝╚══════╝╚═╝  ╚═╝╚═════╝ ╚══════╝

ElevenLabs Integration Guide

Focus: Text-to-speech, voice cloning, speech-to-text, and voice design via ElevenLabs APIs and Claude Code tooling.

Overview

ElevenLabs provides state-of-the-art AI voice generation — synthesize speech, clone voices, transcribe audio, and design custom voices. Used in codeAmani products for voice UI features, audio notifications, and multilingual TTS (including Swahili).

Here is the big picture — text flows out as audio, and user audio flows back in as text, all through one API:

flowchart LR
  A["App text<br/>e.g. dashboard reply"] -->|"text_to_speech"| B["ElevenLabs API"]
  B --> C["Audio stream<br/>mp3"]
  C --> D["Play to user<br/>IVR or voice note"]
  E["User audio<br/>recording"] -->|"speech_to_text<br/>scribe_v2"| B
  B --> F["Transcribed text"]

Official Documentation

Resource URL
API Reference https://elevenlabs.io/docs/api-reference
Node.js SDK https://github.com/elevenlabs/elevenlabs-js
Python SDK https://github.com/elevenlabs/elevenlabs-python
Voice Library https://elevenlabs.io/voice-library
Models Reference https://elevenlabs.io/docs/models

MCP Server Setup

ElevenLabs offers two MCP paths. Prefer the hosted server — there is nothing to install and it authenticates over OAuth (no API key copied into a config file).

Note: there is no @elevenlabs/elevenlabs-mcp npm package (a common mistake — npx will 404). The self-hosted server is the Python package elevenlabs-mcp (PyPI), and its GitHub repo was archived on 2026-08-20 in favor of the hosted server below.

# Streamable-HTTP remote server; completes an OAuth sign-in on first connect
claude mcp add --transport http elevenlabs https://api.elevenlabs.io/v1/mcp

The hosted server focuses on ElevenLabs Agents management (list/create/update agents in your workspace). Revoke access any time from your ElevenLabs account settings.

Self-hosted (local Python server — creative tools)

For the creative toolset (TTS, STT, voice management, sound generation) run the elevenlabs-mcp PyPI package locally via uvx (requires the uv Python tool):

// .mcp.json
{
  "mcpServers": {
    "elevenlabs": {
      "command": "uvx",
      "args": ["elevenlabs-mcp"],
      "env": {
        "ELEVENLABS_API_KEY": "${ELEVENLABS_API_KEY}"
      }
    }
  }
}

MCP Capabilities (local elevenlabs-mcp server)

Tool Description
text_to_speech Convert text to audio using a specified voice
list_voices Browse available and cloned voices
get_voice Get details and settings for a voice
speech_to_text Transcribe audio files
sound_generation Generate sound effects from text

SDK Setup

Node.js / TypeScript

The npm package was renamed — the old bare elevenlabs package is deprecated ("moved to @elevenlabs/elevenlabs-js"). Install the scoped package (current @elevenlabs/elevenlabs-js is v2.64.0, SDK v2):

npm install @elevenlabs/elevenlabs-js
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";

const client = new ElevenLabsClient({
  apiKey: process.env.ELEVENLABS_API_KEY,
});

Python

The Python package keeps the bare elevenlabs name (current v2.64.0):

pip install elevenlabs
from elevenlabs.client import ElevenLabs

client = ElevenLabs(api_key=os.environ["ELEVENLABS_API_KEY"])

Core Patterns

You are about to wire these up — here is how a streaming TTS request travels from your Next.js route to the listener:

sequenceDiagram
  participant Client
  participant Route as "API route<br/>/api/tts"
  participant EL as "ElevenLabs"
  Client->>Route: POST text + voiceId
  Route->>EL: textToSpeech.stream with modelId
  EL-->>Route: audio chunks
  Route-->>Client: audio/mpeg response
  Client->>Client: play audio

Text-to-Speech (Streaming)

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";

const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });

// Stream audio to a file (v2 SDK: textToSpeech.stream, camelCase params)
const audioStream = await client.textToSpeech.stream(
  "JBFqnCBsd6RMkjVDRZzb", // Voice ID (Rachel — default)
  {
    text: "Karibu! Welcome to your codeAmani dashboard.",
    modelId: "eleven_multilingual_v2",
    voiceSettings: {
      stability: 0.5,
      similarityBoost: 0.75,
    },
  }
);

// Write to disk
import { createWriteStream } from "fs";
const writer = createWriteStream("output.mp3");
for await (const chunk of audioStream) {
  writer.write(chunk);
}
writer.end();

Next.js App Router API Route

// app/api/tts/route.ts
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { NextRequest } from "next/server";

const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });

export async function POST(req: NextRequest) {
  const { text, voiceId = "JBFqnCBsd6RMkjVDRZzb" } = await req.json();

  const audioStream = await client.textToSpeech.stream(voiceId, {
    text,
    modelId: "eleven_multilingual_v2",
  });

  // Collect chunks
  const chunks: Buffer[] = [];
  for await (const chunk of audioStream) {
    chunks.push(Buffer.from(chunk));
  }

  return new Response(Buffer.concat(chunks), {
    headers: {
      "Content-Type": "audio/mpeg",
      "Cache-Control": "no-store",
    },
  });
}

List Available Voices

// v2 SDK: voices.search() (GET /v2/voices, paginated). voices.getAll() still
// works as a legacy alias but search() is the current method.
const { voices } = await client.voices.search();

for (const voice of voices) {
  console.log(`${voice.voice_id}: ${voice.name} (${voice.labels?.language ?? "multi"})`);
}

Speech-to-Text (Transcription)

import { createReadStream } from "fs";

const transcription = await client.speechToText.convert({
  file: createReadStream("recording.mp3"),
  modelId: "scribe_v2", // current STT model (scribe_v1 is deprecated)
  languageCode: "sw", // Swahili
});

console.log(transcription.text);

Voice cloning

Instant Voice Cloning (IVC) turns a short audio sample into a reusable voice. You upload one or more recordings, get back a voice_id, then synthesize speech with it like any other voice. This is how you give a codeAmani product its own branded voice — or let a user respond in their own voice for voice-note replies.

The flow is two API calls — clone once, reuse the voice_id forever:

flowchart LR
  A["Audio samples<br/>clean · 1+ min"] -->|"voices.ivc.create"| B["ElevenLabs"]
  B --> C["voice_id<br/>saved to account"]
  C -->|"textToSpeech.convert"| D["Synthesized audio<br/>in cloned voice"]
  C --> E["Store voice_id<br/>in your DB"]

Create a cloned voice, then synthesize with it

import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import fs from "fs";

const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });

// 1. Clone a voice from one or more audio samples
const cloned = await client.voices.ivc.create({
  name: "codeAmani Brand Voice",
  files: [
    fs.createReadStream("samples/sample-1.mp3"),
    fs.createReadStream("samples/sample-2.mp3"),
  ],
});

const voiceId = cloned.voice_id;
console.log("Cloned voice id:", voiceId);
// Persist voiceId in your DB so you can reuse it without re-cloning.

// 2. Synthesize speech using the returned voice_id
const audio = await client.textToSpeech.convert(voiceId, {
  text: "Karibu! Hii ni sauti yako mpya kwenye codeAmani.",
  modelId: "eleven_multilingual_v2",
  outputFormat: "mp3_44100_128",
});

const chunks: Buffer[] = [];
for await (const chunk of audio) {
  chunks.push(Buffer.from(chunk));
}
fs.writeFileSync("cloned-output.mp3", Buffer.concat(chunks));

The create call returns an AddVoiceIVCResponseModel with voice_id (use this everywhere) and requires_verification (whether the voice must be verified before high-volume use). The cloned voice also appears in client.voices.getAll().

Gotcha — consent and sample quality. Only clone voices you have explicit permission to use; cloning a real person's voice without consent breaks ElevenLabs' terms (and KDPA-style consent expectations for biometric/voice data). Quality is bounded by your samples: use clean, single-speaker recordings with no background noise or music — a minute of crisp audio beats ten minutes of noisy phone audio. Keep the API key server-side; never expose it to the client doing the upload.


Model Reference

Model ID Use Case Languages Latency
eleven_v3 Flagship — most expressive/emotional TTS, long-form 70+ higher
eleven_v3_conversational Expressive real-time TTS for voice agents 70+ ~280ms
eleven_multilingual_v2 Lifelike, consistent high-quality TTS 29 ~1s
eleven_flash_v2_5 Lowest latency / cheapest (50% off per char), bulk & real-time 32 ~75ms
scribe_v2 Speech-to-text transcription (batch) 90+ —
scribe_v2_realtime Streaming speech-to-text 90+ ~150ms

Deprecated / superseded (do not use in new code): eleven_turbo_v2_5 and eleven_turbo_v2 — ElevenLabs recommends the Flash models in all use cases (functionally equivalent, lower latency). scribe_v1 → migrate to scribe_v2.

For Swahili TTS, eleven_multilingual_v2 remains the proven choice for quality; eleven_v3 (70+ languages) is the newer, more expressive option and eleven_flash_v2_5 (32 languages) covers low-latency/real-time Swahili.


Webhooks

ElevenLabs sends webhook events for async operations (e.g. batch speech-to-text, post-call transcripts). The signature is HMAC-SHA256 in the ElevenLabs-Signature header, formatted t=<unix_ts>,v0=<hex_hmac>, where the signed message is `${timestamp}.${rawBody}` (not the raw body alone). Reject requests whose timestamp is outside a ~30-minute window to block replays.

Prefer the SDK helper, which verifies the signature, checks the timestamp window, and parses the payload for you:

// app/api/webhooks/elevenlabs/route.ts
import { ElevenLabsClient } from "@elevenlabs/elevenlabs-js";
import { NextRequest } from "next/server";

const client = new ElevenLabsClient({ apiKey: process.env.ELEVENLABS_API_KEY });

export async function POST(req: NextRequest) {
  const body = await req.text();
  const sigHeader = req.headers.get("ElevenLabs-Signature") ?? "";
  const secret = process.env.ELEVENLABS_WEBHOOK_SECRET!;

  let event;
  try {
    event = await client.webhooks.constructEvent(body, sigHeader, secret);
  } catch {
    return new Response("Unauthorized", { status: 401 });
  }

  // Handle event.type, e.g. "speech_to_text_transcription.completed"
  return new Response("OK");
}

If you verify manually, recompute HMAC_SHA256(secret, ${t}.${body}) and compare (constant-time) against the v0= value parsed from the header.


Environment Variables

# Required
ELEVENLABS_API_KEY=sk_...

# Optional
ELEVENLABS_WEBHOOK_SECRET=whsec_...
ELEVENLABS_DEFAULT_VOICE_ID=JBFqnCBsd6RMkjVDRZzb

Common Use Cases

Use Case Approach
Voice notifications POST to /api/tts, play on client
Swahili audio content eleven_multilingual_v2 + languageCode: "sw"
Voice cloning voices.ivc.create() with audio samples → get voice_id → use in TTS (see Voice cloning)
Audio transcription scribe_v2 model + speechToText.convert()
Real-time voice AI WebSocket streaming with eleven_flash_v2_5 (or eleven_v3_conversational)

Troubleshooting

Issue Fix
401 Unauthorized Check ELEVENLABS_API_KEY is set and valid
Audio sounds robotic Increase stability (0.7–0.9) and similarity_boost (0.8)
Swahili not accurate Use eleven_multilingual_v2 (or eleven_v3); Flash trades some quality for latency
Large audio files slow Use streaming (textToSpeech.stream) instead of buffered response
MCP server not connecting Run claude mcp list and check ELEVENLABS_API_KEY in env

Official docs: