xAI Integration Guide

Technology: xai · Category: ai · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/xai

Insight:

xAI is the real-time grounding tier — Grok's server-side web_search and x_search tools give a frontier reasoning model live web + X data with inline citations, no scraping pipeline needed. Flagship grok-4.6 runs a 500K-token context; grok-4.3 stays on for 1M-token long-context jobs. The API is OpenAI-compatible, so it drops into existing OpenAI-SDK or AI-SDK code with a baseURL/provider swap.

██╗  ██╗ █████╗ ██╗
╚██╗██╔╝██╔══██╗██║
 ╚███╔╝ ███████║██║
 ██╔██╗ ██╔══██║██║
██╔╝ ██╗██║  ██║██║
╚═╝  ╚═╝╚═╝  ╚═╝╚═╝

xAI Integration Guide

Focus: Grok models via the xAI API — real-time Live Search (web + X), reasoning, structured outputs, and image generation, through @ai-sdk/xai (TypeScript) or the official xai-sdk (Python).

Overview

xAI serves the Grok model family at https://api.x.ai/v1. The flagship grok-4.6 is a frontier reasoning model with a 500,000-token context window, text + image input, and access to xAI's signature capability: server-side Live Search tools (web_search, x_search) that let the model browse the web and X in real time and return inline citations. When a job genuinely needs more room than 500K tokens, grok-4.3 stays available as the 1M-token long-context option. The API is OpenAI-compatible — the OpenAI SDK works with a base-URL swap — and xAI also ships an official Python SDK plus first-class support in the Vercel AI SDK.

flowchart LR
  A["Your app code"] --> B{"Which client?"}
  B -->|"TypeScript"| C["@ai-sdk/xai<br/>Vercel AI SDK provider"]
  B -->|"Python"| D["xai-sdk<br/>official gRPC client"]
  B -->|"Any language"| E["OpenAI SDK<br/>baseURL · api.x.ai/v1"]
  C --> F["grok-4.6<br/>500K-token context"]
  D --> F
  E --> F
  F --> G["web_search + x_search<br/>server-side Live Search<br/>with citations"]

Official Documentation

Resource URL
Quickstart https://docs.x.ai/developers/quickstart
Models https://docs.x.ai/developers/models/grok-4.6
Web Search tool https://docs.x.ai/developers/tools/web-search
Streaming https://docs.x.ai/developers/model-capabilities/text/streaming
Console (keys, billing) https://console.x.ai
Python SDK https://github.com/xai-org/xai-sdk-python

SDK Setup

xAI has no official JS SDK; its docs use the AI SDK provider.

npm i @ai-sdk/xai ai
import { xai } from "@ai-sdk/xai"; // reads XAI_API_KEY from env
import { generateText } from "ai";

const { text } = await generateText({
  model: xai("grok-4.6"),
  prompt: "Summarize today's Kenyan mobile-money news.",
});

Custom configuration:

import { createXai } from "@ai-sdk/xai";

const xai = createXai({ apiKey: process.env.XAI_API_KEY });

TypeScript — OpenAI SDK (drop-in compatibility)

import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.XAI_API_KEY,
  baseURL: "https://api.x.ai/v1",
  timeout: 360000, // reasoning models can think for a while
});

Python — official xai-sdk

pip install xai-sdk   # Python >= 3.10
from xai_sdk import Client
from xai_sdk.chat import user

client = Client()  # reads XAI_API_KEY from environment
chat = client.chat.create(model="grok-4.6")
chat.append(user("Hello, how are you?"))
response = chat.sample()
print(response.content)

Models

Model ID Aliases Modalities Context Notes
grok-4.6 grok-4.6-latest text + image → text 500K tokens Flagship frontier reasoning model — coding, agentic tasks, knowledge work
grok-4.5 grok-4.5-latest text + image → text 500K tokens Prior-generation reasoning model
grok-4.3 grok-4.3-latest text + image → text 1M tokens Long-context option — reach for it only when >500K tokens is genuinely needed
grok-4.20-0309-reasoning grok-4.20-0309-reasoning-latest text + image → text 1M tokens Reasoning variant (non-reasoning + multi-agent variants also exist)
grok-build-0.1 grok-build-0.1-latest text → text 256K tokens Function calling, structured outputs, reasoning
grok-imagine-image-2.0 — text → image — Image generation (also grok-imagine-image-quality, grok-imagine-image)

Model lineup and pricing move fast — check https://docs.x.ai/developers/models and the console before hardcoding a model ID; prefer -latest aliases only for experiments, pinned IDs in production. grok-4.6 pricing (2026-08): $2/M input ($4/M above 200K tokens), $0.50/M cached, $6/M output ($12/M above 200K).


Core Patterns

Chat Completion (OpenAI-compatible)

curl https://api.x.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-4.6",
    "messages": [
      {"role": "system", "content": "You are Grok, a helpful and truthful AI assistant built by xAI."},
      {"role": "user", "content": "What is prompt caching?"}
    ]
  }'

Streaming

Set a long timeout — reasoning models can run for minutes before the first token.

const stream = await client.chat.completions.create({
  model: "grok-4.6",
  messages: [{ role: "user", content: prompt }],
  stream: true,
});

for await (const chunk of stream) {
  process.stdout.write(chunk.choices[0]?.delta?.content ?? "");
}

Live Search — web + X with citations

The killer feature: server-side search tools. Declare them in tools and the model plans, searches, browses, and cites — no tool-execution loop on your side. Inline citations are on by default when web_search is enabled.

curl https://api.x.ai/v1/responses \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $XAI_API_KEY" \
  -d '{
    "model": "grok-4.6",
    "input": [{"role": "user", "content": "What is the latest update from xAI?"}],
    "tools": [{"type": "web_search"}, {"type": "x_search"}]
  }'
from xai_sdk import Client
from xai_sdk.chat import user
from xai_sdk.tools import web_search, x_search

client = Client()
chat = client.chat.create(
    model="grok-4.6",
    tools=[web_search(), x_search()],
)
chat.append(user("What is the latest update from xAI?"))
response = chat.sample()
print(response.content)
print(response.citations)               # source URLs
print(response.server_side_tool_usage)  # per-tool call counts

Scope searches with domain filters (max 5 domains; allowed_domains and excluded_domains are mutually exclusive):

// AI SDK form
xai.tools.webSearch({ allowedDomains: ["centralbank.go.ke"] });
# xai-sdk form
web_search(allowed_domains=["centralbank.go.ke"])

Structured Outputs (AI SDK + Zod)

import { xai } from "@ai-sdk/xai";
import { generateText, Output } from "ai";
import { z } from "zod";

const InvoiceSchema = z.object({
  vendor_name: z.string().describe("Name of the vendor"),
  invoice_number: z.string().describe("Unique invoice identifier"),
  total_amount: z.number().min(0).describe("Total amount due"),
});

const result = await generateText({
  model: xai.responses("grok-4.6"),
  output: Output.object({ schema: InvoiceSchema }),
  prompt: "Extract the invoice fields from this email: ...",
});

Prompt Caching — x-grok-conv-id

Requests carrying the same conversation ID route to the same server, maximizing cache hits. Check usage.prompt_tokens_details.cached_tokens to verify.

const response = await client.chat.completions.create(
  { model: "grok-4.6", messages },
  { headers: { "x-grok-conv-id": conversationId } },
);
console.log(response.usage?.prompt_tokens_details?.cached_tokens);

Image Generation

import { xai } from "@ai-sdk/xai";
import { experimental_generateImage as generateImage } from "ai";

const { image } = await generateImage({
  model: xai.image("grok-imagine-image-quality"),
  prompt: "A boda boda rider at sunset in Nairobi",
});

AI Routing: When to Use xAI

In codeAmani's routing policy, Anthropic Claude stays primary for complex reasoning and code gen. xAI earns its slot when the answer needs live web or X data:

flowchart TD
  A["Incoming task"] --> B{"Needs real-time<br/>web or X data?"}
  B -->|"yes"| C["xAI grok-4.6<br/>web_search + x_search"]
  B -->|"no"| D{"Complex reasoning<br/>or code gen?"}
  D -->|"yes"| E["Anthropic Claude"]
  D -->|"no"| F["Route by cost tier<br/>(DeepSeek / Haiku)"]
Task Recommended
Real-time news / market / social sentiment grok-4.6 + Live Search
"What's happening on X about …" grok-4.6 + x_search
Ground a >500K-token corpus in live data grok-4.3 (1M context) + Live Search
Complex reasoning, agents, code gen Anthropic Claude
Bulk cheap inference DeepSeek / Haiku tier

Environment Variables

# Required — create at https://console.x.ai
XAI_API_KEY=xai-...

codeAmani notes


Troubleshooting

Issue Fix
401 Unauthorized Verify XAI_API_KEY (starts with xai-) and that it's set server-side
Request times out Reasoning models think before answering — raise SDK timeout (docs use up to 3600s)
allowed_domains rejected Max 5 domains; cannot combine with excluded_domains in one request
No citations in response Citations require a search tool (web_search / x_search) in tools
cached_tokens always 0 Send a stable x-grok-conv-id header so requests hit the same cache
Model not found Check current IDs at docs.x.ai/developers/models — lineup rotates quickly

Official docs: