← Back to dashboard

DeepSeek Integration Guide

What is DeepSeek?

The real model

Mixture-of-Experts models, OpenAI-compat API, MIT-licensed weights.

V4 Pro is a ~1.6T-parameter MoE with ~49B active per token; V4 Flash is ~284B with ~13B active. Both carry a 1M-token context and a dual thinking mode — enable it per request with `reasoning_effort` + `thinking: { type: "enabled" }`, and the model emits `reasoning_content` before the final answer. The hosted API uses the same `/chat/completions` schema as OpenAI — point your SDK at `https://api.deepseek.com` and most code Just Works. Function calling is supported. Weights are on Hugging Face under MIT — the efficient FP4/FP8 Flash build is ~146 GB (Blackwell-class GPUs). For codeAmani, DeepSeek is the cost-optimisation lever: route the cheapest ~80% of traffic to `deepseek-v4-flash`, fall back to Claude when quality matters.

Five reasons DeepSeek is worth integrating

Open weights + cheap API + OpenAI-compatible schema = unusually low switching cost.

Text
██████╗ ███████╗███████╗██████╗ ███████╗███████╗███████╗██╗  ██╗
██╔══██╗██╔════╝██╔════╝██╔══██╗██╔════╝██╔════╝██╔════╝██║ ██╔╝
██║  ██║█████╗  █████╗  ██████╔╝███████╗█████╗  █████╗  █████╔╝
██║  ██║██╔══╝  ██╔══╝  ██╔═══╝ ╚════██║██╔══╝  ██╔══╝  ██╔═██╗
██████╔╝███████╗███████╗██║     ███████║███████╗███████╗██║  ██╗
╚═════╝ ╚══════╝╚══════╝╚═╝     ╚══════╝╚══════╝╚══════╝╚═╝  ╚═╝

DeepSeek Integration Guide

Focus: Cost-effective AI inference and chain-of-thought reasoning — the DeepSeek-V4 family (deepseek-v4-flash / deepseek-v4-pro) via the OpenAI-compatible SDK.

Overview

DeepSeek's current generation is the V4 family, served under two model IDs: deepseek-v4-flash (fast, very cheap — the default cost tier) and deepseek-v4-pro (higher quality). Both are OpenAI-compatible — swap the base URL and API key, keep the same code — carry a 1M-token context window, and support a dual thinking / non-thinking mode: the old V3-chat / R1-reasoner split is gone, and chain-of-thought is now a per-request toggle on the same model. An experimental multimodal variant, deepseek-v4-flash-vision-exp, adds image input. Used in codeAmani products as a cost-optimization alternative for tasks that don't require Anthropic's highest capability tier.

Migration note: the legacy deepseek-chat (V3) and deepseek-reasoner (R1) IDs were retired after 2026-07-24. Move existing calls to deepseek-v4-flash (drop-in replacement for deepseek-chat) or deepseek-v4-pro with thinking enabled (replacement for deepseek-reasoner).

Here's the big picture — the same OpenAI SDK call points at DeepSeek, picks a tier, and toggles thinking per request:

Official Documentation


SDK Setup

DeepSeek is OpenAI API-compatible — use the official OpenAI SDK with a custom base URL.

Bash
npm install openai
TypeScript
import OpenAI from "openai";

const deepseek = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY!,
  baseURL: "https://api.deepseek.com",
});

Models

Model IDServesBest ForContextThinking mode
deepseek-v4-flashDeepSeek-V4-Flash-0731Fast chat, code, structured output, high-volume bulk1M tokensoptional (per request)
deepseek-v4-proDeepSeek-V4-Pro-0813Higher-quality answers, harder reasoning/math/logic1M tokensoptional (per request)
deepseek-v4-flash-vision-expexperimental multimodalImage input; matches Flash on text1M tokensoptional (per request)

Both deepseek-v4-flash and deepseek-v4-pro support the dual thinking / non-thinking mode — reasoning is a per-request toggle, not a separate model (see Reasoning (R1-style) — thinking mode below). Use exact IDs; DeepSeek rolls new checkpoints (e.g. -0731, -0813) under the stable base ID, so keep using deepseek-v4-flash / deepseek-v4-pro.

Retired: deepseek-chat and deepseek-reasoner (the V3/R1 IDs) were retired after 2026-07-24 — do not use them in new code.


Core Patterns

Standard Chat Completion

TypeScript
const response = await deepseek.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    { role: "system", content: "You are a helpful assistant for codeAmani Labs." },
    { role: "user", content: "Summarize this M-Pesa transaction log." },
  ],
  max_tokens: 1024,
});

console.log(response.choices[0].message.content);

Reasoning (R1-style) — thinking mode

Chain-of-thought is now a per-request thinking toggle on the V4 models — enable it with reasoning_effort plus DeepSeek's thinking extension. When enabled, the model exposes its reasoning in reasoning_content before the final content.

TypeScript
const response = await deepseek.chat.completions.create({
  model: "deepseek-v4-pro",
  messages: [
    { role: "user", content: "Why is my Supabase RLS policy blocking authenticated users?" },
  ],
  reasoning_effort: "high",
  // `thinking` is a DeepSeek extension not in the OpenAI types; the SDK forwards it.
  // @ts-expect-error — deepseek-specific field
  thinking: { type: "enabled" },
  max_tokens: 4096,
});

const choice = response.choices[0];
// @ts-expect-error — deepseek-specific field
console.log("Reasoning:", choice.message.reasoning_content);
console.log("Answer:", choice.message.content);

Do not feed reasoning_content back into message history — it is intermediate scratch-work, not part of the conversation.

Streaming Response

TypeScript
const stream = await deepseek.chat.completions.create({
  model: "deepseek-v4-flash",
  stream: true,
  messages: [{ role: "user", content: prompt }],
});

for await (const chunk of stream) {
  const delta = chunk.choices[0]?.delta?.content ?? "";
  process.stdout.write(delta);
}

Next.js App Router Streaming Route

TypeScript
// app/api/ai/deepseek/route.ts
import OpenAI from "openai";
import { NextRequest } from "next/server";

const deepseek = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY!,
  baseURL: "https://api.deepseek.com",
});

export async function POST(req: NextRequest) {
  const { messages } = await req.json();

  const stream = await deepseek.chat.completions.create({
    model: "deepseek-v4-flash",
    stream: true,
    messages,
  });

  const encoder = new TextEncoder();
  const readable = new ReadableStream({
    async start(controller) {
      for await (const chunk of stream) {
        const text = chunk.choices[0]?.delta?.content ?? "";
        if (text) controller.enqueue(encoder.encode(text));
      }
      controller.close();
    },
  });

  return new Response(readable, {
    headers: { "Content-Type": "text/plain; charset=utf-8" },
  });
}

AI Routing: When to Use DeepSeek

In codeAmani's AI routing strategy, DeepSeek slots in as a cost-optimization tier:

This decision flow shows exactly where DeepSeek earns its place alongside Anthropic — pick the right tier and you save cost without losing quality:

TypeScript
// lib/ai.ts
type TaskComplexity = "high" | "medium" | "low";

function selectModel(complexity: TaskComplexity, requiresReasoning: boolean) {
  if (complexity === "high" || requiresReasoning) {
    return { provider: "anthropic", model: "claude-sonnet-4-6" };
  }
  if (complexity === "medium") {
    return { provider: "deepseek", model: "deepseek-v4-flash" };
  }
  // Low complexity: fast classification / simple Q&A
  return { provider: "anthropic", model: "claude-haiku-4-5-20251001" };
}
TaskRecommended Model
Complex reasoning, agentsclaude-sonnet-4-6
Step-by-step math / logicdeepseek-v4-pro (thinking on)
Standard Q&A, summariesdeepseek-v4-flash
Fast classificationclaude-haiku-4-5-20251001
Structured JSON outputgpt-4o

Error handling, retries & fallback

The routing section above treats DeepSeek as a cost tier, not a hard dependency — so any call that hits DeepSeek must be able to fall back to Anthropic Claude when DeepSeek throttles or errors. DeepSeek's own docs explicitly suggest this: on a 429, they recommend you "temporarily switch to alternative LLM providers."

Documented status codes

These are the status codes DeepSeek documents on its error codes page (unchanged under V4). Treat the transient ones as retry-then-fallback, and the terminal ones as fail-fast (retrying won't help):

CodeMeaningClassAction
400Invalid request body formatterminalFix the request — do not retry
401Authentication fails (wrong API key)terminalFix DEEPSEEK_API_KEY
402Insufficient balanceterminalTop up; fall back immediately
422Invalid parametersterminalFix params — do not retry
429Rate limit reached (concurrency limit)transientBack off, then fall back
500Server errortransientRetry after a brief wait
503Server overloaded (high traffic)transientRetry after a brief wait

DeepSeek does not publish a fixed requests-per-second limit. Instead it documents a per-user_id concurrency limit (rate limit docs); exceeding the number of in-flight connections is what returns 429. The docs give no prescribed backoff schedule, so the pattern below uses standard exponential backoff with jitter.

Try DeepSeek with backoff, then fall back to Claude

This mirrors the router in lib/ai.ts — selectModel chooses the tier, this wrapper makes the DeepSeek tier resilient. Terminal errors (4xx except 429) skip retries and fall straight through to Claude.

TypeScript
// lib/ai-resilient.ts
import OpenAI from "openai";
import Anthropic from "@anthropic-ai/sdk";

const deepseek = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY!,
  baseURL: "https://api.deepseek.com",
});

const anthropic = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY! });

// Transient per DeepSeek docs: 429 (rate limit), 500 (server error), 503 (overloaded).
const RETRYABLE = new Set([429, 500, 503]);
const MAX_RETRIES = 3;

const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

/**
 * Run a prompt through DeepSeek with exponential backoff on transient errors,
 * then fall back to Anthropic Claude if DeepSeek is unavailable or non-retryable.
 */
export async function completeWithFallback(prompt: string): Promise<string> {
  for (let attempt = 0; attempt <= MAX_RETRIES; attempt++) {
    try {
      const res = await deepseek.chat.completions.create({
        model: "deepseek-v4-flash",
        messages: [{ role: "user", content: prompt }],
        max_tokens: 2048,
      });
      return res.choices[0]?.message.content ?? "";
    } catch (err) {
      // OpenAI SDK surfaces the HTTP status on err.status
      const status = (err as { status?: number }).status;

      // Non-retryable (400/401/402/422) or out of attempts → break to fallback.
      if (!status || !RETRYABLE.has(status) || attempt === MAX_RETRIES) break;

      // Exponential backoff with full jitter: ~0.5s, 1s, 2s (+ jitter).
      const base = 500 * 2 ** attempt;
      await sleep(base + Math.random() * base);
    }
  }

  // Fallback tier — Anthropic Claude (consistent with lib/ai.ts router).
  const msg = await anthropic.messages.create({
    model: "claude-haiku-4-5-20251001",
    max_tokens: 2048,
    messages: [{ role: "user", content: prompt }],
  });
  const block = msg.content[0];
  return block.type === "text" ? block.text : "";
}

Gotcha — empty lines are not errors. While a request waits to be scheduled, DeepSeek keeps the TCP connection alive by sending empty lines (non-streaming) or : keep-alive SSE comments (streaming) rather than data. The OpenAI SDK handles these for you, but if you parse the raw HTTP/SSE stream yourself, skip those blank/comment lines — do not treat them as a malformed response or trip your retry logic on them. Connections also close after ~10 minutes if inference never starts, so set a client timeout below that and let the fallback path catch it.


JSON / Structured Output

TypeScript
const response = await deepseek.chat.completions.create({
  model: "deepseek-v4-flash",
  response_format: { type: "json_object" },
  messages: [
    {
      role: "system",
      content: "Respond only with valid JSON.",
    },
    {
      role: "user",
      content: "Extract: name, amount, phone from this SMS: 'Confirmed. Ksh500 sent to 0712345678 on 12/5/26'",
    },
  ],
});

const data = JSON.parse(response.choices[0].message.content ?? "{}");
// { name: null, amount: 500, phone: "0712345678" }

Environment Variables

Bash
# Required
DEEPSEEK_API_KEY=sk-...

# No separate base URL needed — set in code: https://api.deepseek.com

Cost Reference

DeepSeek is significantly cheaper than GPT-4o / Claude for many tasks. Rates change frequently — check current pricing at https://api-docs.deepseek.com/quick_start/pricing. Prices below are per 1M tokens, USD, and split into off-peak / peak: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri; every other hour is off-peak (half the peak rate), so most of the day bills at the cheaper column.

ModelInput (cache hit)Input (cache miss)Output
deepseek-v4-flash$0.007 / $0.014$0.22 / $0.44$0.66 / $1.32
deepseek-v4-pro$0.022 / $0.044$0.66 / $1.32$1.98 / $3.96
deepseek-v4-flash-vision-exp$0.007 / $0.014$0.22 / $0.44$0.66 / $1.32

Thinking-mode reasoning tokens are billed at the normal output rate. Cache-hit input is ~30× cheaper than cache-miss — DeepSeek caches prompt prefixes automatically (no cache-control header needed). See reference/pricing-snapshot.md.


Troubleshooting

IssueFix
Authentication failsVerify DEEPSEEK_API_KEY — starts with sk-
model not foundUse exact V4 IDs: deepseek-v4-flash or deepseek-v4-pro (the old deepseek-chat / deepseek-reasoner were retired 2026-07-24)
reasoning_content undefinedOnly populated when thinking mode is enabled (reasoning_effort + thinking: { type: "enabled" })
Streaming stops mid-responseCheck max_tokens — default is low; increase to 4096+
TypeScript errors on reasoning_content / thinkingUse // @ts-expect-error — DeepSeek fields not in OpenAI types