DeepSeek Integration Guide

Technology: deepseek · Category: ai · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/deepseek

Insight:

DeepSeek is the budget reasoning tier — its OpenAI-compatible API drops into existing OpenAI / AI-SDK code with just a baseURL + model change. The V4 family (deepseek-v4-flash / deepseek-v4-pro) folds chat and chain-of-thought into a single model with a per-request thinking toggle and a 1M-token context, at a fraction of frontier-model cost. Use deepseek-v4-flash as a cheap fallback for high-volume, cost-sensitive SME workloads where frontier quality isn't required.

██████╗ ███████╗███████╗██████╗ ███████╗███████╗███████╗██╗  ██╗
██╔══██╗██╔════╝██╔════╝██╔══██╗██╔════╝██╔════╝██╔════╝██║ ██╔╝
██║  ██║█████╗  █████╗  ██████╔╝███████╗█████╗  █████╗  █████╔╝
██║  ██║██╔══╝  ██╔══╝  ██╔═══╝ ╚════██║██╔══╝  ██╔══╝  ██╔═██╗
██████╔╝███████╗███████╗██║     ███████║███████╗███████╗██║  ██╗
╚═════╝ ╚══════╝╚══════╝╚═╝     ╚══════╝╚══════╝╚══════╝╚═╝  ╚═╝

DeepSeek Integration Guide

Focus: Cost-effective AI inference and chain-of-thought reasoning — the DeepSeek-V4 family (deepseek-v4-flash / deepseek-v4-pro) via the OpenAI-compatible SDK.

Overview

DeepSeek's current generation is the V4 family, served under two model IDs: deepseek-v4-flash (fast, very cheap — the default cost tier) and deepseek-v4-pro (higher quality). Both are OpenAI-compatible — swap the base URL and API key, keep the same code — carry a 1M-token context window, and support a dual thinking / non-thinking mode: the old V3-chat / R1-reasoner split is gone, and chain-of-thought is now a per-request toggle on the same model. An experimental multimodal variant, deepseek-v4-flash-vision-exp, adds image input. Used in codeAmani products as a cost-optimization alternative for tasks that don't require Anthropic's highest capability tier.

Migration note: the legacy deepseek-chat (V3) and deepseek-reasoner (R1) IDs were retired after 2026-07-24. Move existing calls to deepseek-v4-flash (drop-in replacement for deepseek-chat) or deepseek-v4-pro with thinking enabled (replacement for deepseek-reasoner).

Here's the big picture — the same OpenAI SDK call points at DeepSeek, picks a tier, and toggles thinking per request:

flowchart LR
  A["Your app code"] --> B["OpenAI SDK<br/>baseURL · api.deepseek.com"]
  B --> C{"Which tier?"}
  C -->|"deepseek-v4-flash"| D["V4 Flash<br/>fast + cheapest"]
  C -->|"deepseek-v4-pro"| E["V4 Pro<br/>higher quality"]
  D --> F{"thinking<br/>enabled?"}
  E --> F
  F -->|"no"| G["Response content"]
  F -->|"yes"| H["reasoning_content<br/>plus answer content"]

Official Documentation

Resource URL
API Docs https://api-docs.deepseek.com/
API Reference https://api-docs.deepseek.com/api/create-chat-completion
Models & Pricing https://api-docs.deepseek.com/quick_start/pricing
OpenAI Compatibility https://api-docs.deepseek.com/quick_start/compatibility_guide
Reasoning (thinking mode) https://api-docs.deepseek.com/guides/reasoning_model

SDK Setup

DeepSeek is OpenAI API-compatible — use the official OpenAI SDK with a custom base URL.

npm install openai
import OpenAI from "openai";

const deepseek = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY!,
  baseURL: "https://api.deepseek.com",
});

Models

Model ID Serves Best For Context Thinking mode
deepseek-v4-flash DeepSeek-V4-Flash-0731 Fast chat, code, structured output, high-volume bulk 1M tokens optional (per request)
deepseek-v4-pro DeepSeek-V4-Pro-0813 Higher-quality answers, harder reasoning/math/logic 1M tokens optional (per request)
deepseek-v4-flash-vision-exp experimental multimodal Image input; matches Flash on text 1M tokens optional (per request)

Both deepseek-v4-flash and deepseek-v4-pro support the dual thinking / non-thinking mode — reasoning is a per-request toggle, not a separate model (see Reasoning (R1-style) — thinking mode below). Use exact IDs; DeepSeek rolls new checkpoints (e.g. -0731, -0813) under the stable base ID, so keep using deepseek-v4-flash / deepseek-v4-pro.

Retired: deepseek-chat and deepseek-reasoner (the V3/R1 IDs) were retired after 2026-07-24 — do not use them in new code.


Core Patterns

Standard Chat Completion

const response = await deepseek.chat.completions.create({
  model: "deepseek-v4-flash",
  messages: [
    { role: "system", content: "You are a helpful assistant for codeAmani Labs." },
    { role: "user", content: "Summarize this M-Pesa transaction log." },
  ],
  max_tokens: 1024,
});

console.log(response.choices[0].message.content);

Reasoning (R1-style) — thinking mode

Chain-of-thought is now a per-request thinking toggle on the V4 models — enable it with reasoning_effort plus DeepSeek's thinking extension. When enabled, the model exposes its reasoning in reasoning_content before the final content.

const response = await deepseek.chat.completions.create({
  model: "deepseek-v4-pro",
  messages: [
    { role: "user", content: "Why is my Supabase RLS policy blocking authenticated users?" },
  ],
  reasoning_effort: "high",
  // `thinking` is a DeepSeek extension not in the OpenAI types; the SDK forwards it.
  // @ts-expect-error — deepseek-specific field
  thinking: { type: "enabled" },
  max_tokens: 4096,
});

const choice = response.choices[0];
// @ts-expect-error — deepseek-specific field
console.log("Reasoning:", choice.message.reasoning_content);
console.log("Answer:", choice.message.content);

Do not feed reasoning_content back into message history — it is intermediate scratch-work, not part of the conversation.

Streaming Response

const stream = await deepseek.chat.completions.create({
  model: "deepseek-v4-flash",
  stream: true,
  messages: [{ role: "user", content: prompt }],
});

for await (const chunk of stream) {
  const delta = chunk.choices[0]?.delta?.content ?? "";
  process.stdout.write(delta);
}

Next.js App Router Streaming Route

// app/api/ai/deepseek/route.ts
import OpenAI from "openai";
import { NextRequest } from "next/server";

const deepseek = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY!,
  baseURL: "https://api.deepseek.com",
});

export async function POST(req: NextRequest) {
  const { messages } = await req.json();

  const stream = await deepseek.chat.completions.create({
    model: "deepseek-v4-flash",
    stream: true,
    messages,
  });

  const encoder = new TextEncoder();
  const readable = new ReadableStream({
    async start(controller) {
      for await (const chunk of stream) {
        const text = chunk.choices[0]?.delta?.content ?? "";
        if (text) controller.enqueue(encoder.encode(text));
      }
      controller.close();
    },
  });

  return new Response(readable, {
    headers: { "Content-Type": "text/plain; charset=utf-8" },
  });
}

AI Routing: When to Use DeepSeek

In codeAmani's AI routing strategy, DeepSeek slots in as a cost-optimization tier:

This decision flow shows exactly where DeepSeek earns its place alongside Anthropic — pick the right tier and you save cost without losing quality:

flowchart TD
  A["Incoming task"] --> B{"High complexity<br/>or needs reasoning?"}
  B -->|"yes"| C["Anthropic<br/>claude-sonnet-4-6"]
  B -->|"no"| D{"Medium complexity?"}
  D -->|"yes"| E["DeepSeek<br/>deepseek-v4-flash"]
  D -->|"no"| F["Anthropic<br/>claude-haiku-4-5"]
// lib/ai.ts
type TaskComplexity = "high" | "medium" | "low";

function selectModel(complexity: TaskComplexity, requiresReasoning: boolean) {
  if (complexity === "high" || requiresReasoning) {
    return { provider: "anthropic", model: "claude-sonnet-4-6" };
  }
  if (complexity === "medium") {
    return { provider: "deepseek", model: "deepseek-v4-flash" };
  }
  // Low complexity: fast classification / simple Q&A
  return { provider: "anthropic", model: "claude-haiku-4-5-20251001" };
}
Task Recommended Model
Complex reasoning, agents claude-sonnet-4-6
Step-by-step math / logic deepseek-v4-pro (thinking on)
Standard Q&A, summaries deepseek-v4-flash
Fast classification claude-haiku-4-5-20251001
Structured JSON output gpt-4o

Error handling, retries & fallback

The routing section above treats DeepSeek as a cost tier, not a hard dependency — so any call that hits DeepSeek must be able to fall back to Anthropic Claude when DeepSeek throttles or errors. DeepSeek's own docs explicitly suggest this: on a 429, they recommend you "temporarily switch to alternative LLM providers."

Documented status codes

These are the status codes DeepSeek documents on its error codes page (unchanged under V4). Treat the transient ones as retry-then-fallback, and the terminal ones as fail-fast (retrying won't help):

Code Meaning Class Action
400 Invalid request body format terminal Fix the request — do not retry
401 Authentication fails (wrong API key) terminal Fix DEEPSEEK_API_KEY
402 Insufficient balance terminal Top up; fall back immediately
422 Invalid parameters terminal Fix params — do not retry
429 Rate limit reached (concurrency limit) transient Back off, then fall back
500 Server error transient Retry after a brief wait
503 Server overloaded (high traffic) transient Retry after a brief wait

DeepSeek does not publish a fixed requests-per-second limit. Instead it documents a per-user_id concurrency limit (rate limit docs); exceeding the number of in-flight connections is what returns 429. The docs give no prescribed backoff schedule, so the pattern below uses standard exponential backoff with jitter.

Try DeepSeek with backoff, then fall back to Claude

This mirrors the router in lib/ai.ts — selectModel chooses the tier, this wrapper makes the DeepSeek tier resilient. Terminal errors (4xx except 429) skip retries and fall straight through to Claude.

// lib/ai-resilient.ts
import OpenAI from "openai";
import Anthropic from "@anthropic-ai/sdk";

const deepseek = new OpenAI({
  apiKey: process.env.DEEPSEEK_API_KEY!,
  baseURL: "https://api.deepseek.com",
});

const anthropic = new Anthropic({ apiKey: process.env.ANTHROPIC_API_KEY! });

// Transient per DeepSeek docs: 429 (rate limit), 500 (server error), 503 (overloaded).
const RETRYABLE = new Set([429, 500, 503]);
const MAX_RETRIES = 3;

const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));

/**
 * Run a prompt through DeepSeek with exponential backoff on transient errors,
 * then fall back to Anthropic Claude if DeepSeek is unavailable or non-retryable.
 */
export async function completeWithFallback(prompt: string): Promise<string> {
  for (let attempt = 0; attempt <= MAX_RETRIES; attempt++) {
    try {
      const res = await deepseek.chat.completions.create({
        model: "deepseek-v4-flash",
        messages: [{ role: "user", content: prompt }],
        max_tokens: 2048,
      });
      return res.choices[0]?.message.content ?? "";
    } catch (err) {
      // OpenAI SDK surfaces the HTTP status on err.status
      const status = (err as { status?: number }).status;

      // Non-retryable (400/401/402/422) or out of attempts → break to fallback.
      if (!status || !RETRYABLE.has(status) || attempt === MAX_RETRIES) break;

      // Exponential backoff with full jitter: ~0.5s, 1s, 2s (+ jitter).
      const base = 500 * 2 ** attempt;
      await sleep(base + Math.random() * base);
    }
  }

  // Fallback tier — Anthropic Claude (consistent with lib/ai.ts router).
  const msg = await anthropic.messages.create({
    model: "claude-haiku-4-5-20251001",
    max_tokens: 2048,
    messages: [{ role: "user", content: prompt }],
  });
  const block = msg.content[0];
  return block.type === "text" ? block.text : "";
}
flowchart TD
  A["completeWithFallback"] --> B["Call DeepSeek<br/>deepseek-v4-flash"]
  B --> C{"Result?"}
  C -->|"success"| D["Return content"]
  C -->|"429 · 500 · 503"| E{"Retries left?"}
  E -->|"yes"| F["Backoff with jitter<br/>then retry"]
  F --> B
  E -->|"no"| G["Fallback to Claude<br/>claude-haiku-4-5"]
  C -->|"400 · 401 · 402 · 422"| G
  G --> D

Gotcha — empty lines are not errors. While a request waits to be scheduled, DeepSeek keeps the TCP connection alive by sending empty lines (non-streaming) or : keep-alive SSE comments (streaming) rather than data. The OpenAI SDK handles these for you, but if you parse the raw HTTP/SSE stream yourself, skip those blank/comment lines — do not treat them as a malformed response or trip your retry logic on them. Connections also close after ~10 minutes if inference never starts, so set a client timeout below that and let the fallback path catch it.


JSON / Structured Output

const response = await deepseek.chat.completions.create({
  model: "deepseek-v4-flash",
  response_format: { type: "json_object" },
  messages: [
    {
      role: "system",
      content: "Respond only with valid JSON.",
    },
    {
      role: "user",
      content: "Extract: name, amount, phone from this SMS: 'Confirmed. Ksh500 sent to 0712345678 on 12/5/26'",
    },
  ],
});

const data = JSON.parse(response.choices[0].message.content ?? "{}");
// { name: null, amount: 500, phone: "0712345678" }

Environment Variables

# Required
DEEPSEEK_API_KEY=sk-...

# No separate base URL needed — set in code: https://api.deepseek.com

Cost Reference

DeepSeek is significantly cheaper than GPT-4o / Claude for many tasks. Rates change frequently — check current pricing at https://api-docs.deepseek.com/quick_start/pricing. Prices below are per 1M tokens, USD, and split into off-peak / peak: peak hours are 01:00–04:00 and 06:00–10:00 UTC, Mon–Fri; every other hour is off-peak (half the peak rate), so most of the day bills at the cheaper column.

Model Input (cache hit) Input (cache miss) Output
deepseek-v4-flash $0.007 / $0.014 $0.22 / $0.44 $0.66 / $1.32
deepseek-v4-pro $0.022 / $0.044 $0.66 / $1.32 $1.98 / $3.96
deepseek-v4-flash-vision-exp $0.007 / $0.014 $0.22 / $0.44 $0.66 / $1.32

Thinking-mode reasoning tokens are billed at the normal output rate. Cache-hit input is ~30× cheaper than cache-miss — DeepSeek caches prompt prefixes automatically (no cache-control header needed). See reference/pricing-snapshot.md.


Troubleshooting

Issue Fix
Authentication fails Verify DEEPSEEK_API_KEY — starts with sk-
model not found Use exact V4 IDs: deepseek-v4-flash or deepseek-v4-pro (the old deepseek-chat / deepseek-reasoner were retired 2026-07-24)
reasoning_content undefined Only populated when thinking mode is enabled (reasoning_effort + thinking: { type: "enabled" })
Streaming stops mid-response Check max_tokens — default is low; increase to 4096+
TypeScript errors on reasoning_content / thinking Use // @ts-expect-error — DeepSeek fields not in OpenAI types

Official docs: