Grok Bot Developer Guide

AUDITED · SEP 1 2026 · docs.x.ai

6 stages · 3 surfaces · live cost model

Grok Bot
Developer Portal

A bot is not an API call — it is a stateful loop wrapped around a stateless model. The model is one of six jobs; these are the other five, and what each one costs.

One turn, end to end

The Bot Loop

What runs~1 ms

Recompute the HMAC over the RAW request body and compare in constant time. Read req.text() first — parsing to JSON before verifying changes the bytes and the signature will never match.

What breaks without it

Skip it and anyone who learns your webhook URL can post messages as any user. This is the only thing standing between your bot and an open relay.

route.ts
// app/api/whatsapp/webhook/route.ts
import { after } from "next/server";
import { handleTurn } from "@/lib/bot/turn";

export async function POST(req: Request) {
  // Signature is computed over the RAW body - read text, never req.json() first
  const raw = await req.text();
  if (!verifySignature(raw, req.headers.get("x-hub-signature-256"))) {
    return new Response("invalid signature", { status: 401 });
  }

  const payload = JSON.parse(raw);
  const msg = payload.entry?.[0]?.changes?.[0]?.value?.messages?.[0];

  // ACK first, work after - Meta retries anything it does not get a fast 200 on
  if (msg) after(() => handleTurn(msg));
  return new Response("ok", { status: 200 });
}
WhatsApp-first

Chat Surfaces

WhatsApp Cloud API

GA
Transport
REST over Graph API (no SDK)
Streaming
None — one message per send
Message cap
~4096 chars
Webhook
Fast 200 or Meta retries
The 24-hour window

Outside 24h from the user's last message, only pre-approved template messages are delivered. Free-form replies are rejected.

Telegram

Transport
Bot API via grammy
Streaming
Simulated (editMessageText)
Message cap
4096 chars
Webhook
Standard webhook
429 with retry_after

Telegram tells you exactly how long to wait in retry_after — honour it rather than applying your own backoff.

Discord

Transport
Gateway + REST via discord.js
Streaming
Simulated (message edits)
Message cap
2000 chars
Webhook
3-second interaction deadline
deferReply() or fail

You must acknowledge an interaction within 3 seconds. Any model call is slower than that, so deferReply() is mandatory, not optional.

Setup

Getting Started

install
npm install ai @ai-sdk/xai @ai-sdk/anthropic zod
npm install grammy          # Telegram
# WhatsApp Cloud API is plain REST - no SDK needed

Every one of these is server-side only. An XAI_API_KEY in a NEXT_PUBLIC_* variable is shipped to every browser that loads the page.

Window + summary

Conversation State

Keep the last N turns verbatim and compress everything older into one paragraph the persona can read. The window is the single highest-leverage cost control you have — the calculator below prices it.

history.ts
// lib/bot/history.ts
import type { ModelMessage } from "ai";

const VERBATIM_TURNS = 12;

export async function buildMessages(conversationId: string): Promise<ModelMessage[]> {
  const convo = await getConversation(conversationId);
  const recent = await getRecentMessages(conversationId, VERBATIM_TURNS);

  const messages: ModelMessage[] = [];
  if (convo.summary) {
    messages.push({
      role: "user",
      content: `[Earlier in this conversation]\n${convo.summary}`,
    });
  }
  for (const m of recent) {
    messages.push({ role: m.role, content: m.content } as ModelMessage);
  }
  return messages;
}
Four layers, in order

Persona Design

1

Scope fence

What the bot is for, and explicitly what it will not do. First, because it is the instruction most likely to survive a long transcript.

2

Voice

Tone, length, language. For codeAmani bots this is where Swahili / Sheng register and code-switching rules live.

3

Hard rules

Never invent prices. Never promise delivery dates. Never repeat a customer's full phone number back to them.

4

Escalation

The exact conditions for calling handoff_to_human — anger, refunds, anything involving money moving.

Version the persona like code. A prompt change is a behaviour change, and you will want to know which release started apologising in the wrong language.

The model asks, your code decides

Tool Loop

Scope every tool to the caller

Take chat_id from experimental_context, never from a model-supplied argument. The model will happily pass someone else's id.

Validate with a schema, not trust

inputSchema is a boundary. Treat tool arguments exactly like untrusted form input, because that is what they are.

Cap the steps

stopWhen: stepCountIs(5). Each step is a full prompt resend at full price.

tools.ts
// lib/bot/tools.ts
import { tool } from "ai";
import { z } from "zod";

export const botTools = {
  checkOrder: tool({
    description: "Look up an order by its reference for the CURRENT customer.",
    inputSchema: z.object({ reference: z.string().min(4) }),
    // The model asks. Your code decides - and scopes to the caller.
    execute: async ({ reference }, { experimental_context }) => {
      const { chatId } = experimental_context as { chatId: string };
      const order = await db.order.findFirst({ where: { reference, chatId } });
      return order ?? { error: "not_found" };
    },
  }),

  handoffToHuman: tool({
    description: "Escalate to a human agent when the customer asks or is upset.",
    inputSchema: z.object({ reason: z.string() }),
    execute: async ({ reason }, { experimental_context }) => {
      const { chatId } = experimental_context as { chatId: string };
      await notifyAgents({ chatId, reason });
      return { escalated: true };
    },
  }),
};
Live model · verified Sep 1, 2026

Cost per Conversation

Tokens per request is the wrong unit for a bot. The unit is cost per conversation, because that is what scales with users. Drag the sliders to price your own bot.

ModelContextTypeIn / OutCached inWhy you would pick it
Grok 4.6GA
grok-4.6
500kFast$2 / $6$0.5xAI's current general-purpose recommendation and the sensible default bot brain.
Grok 4.5GA
grok-4.5
500kFast$2 / $6$0.3Previous generation. Same headline price as 4.6 with a cheaper cached-input rate.
Grok 4.3GA
grok-4.3
1MFast$1.25 / $2.5$0.2Cheapest per token with the largest context. The value pick for high-volume bots.
Grok 4.20 ReasoningGA
grok-4.20-0309-reasoning
1MReasoning$1.25 / $2.5$0.2Reasoning variant. Same token price as 4.3 but latency is the tax — often too slow for a chat surface.
Grok Build 0.1GA
grok-build-0.1
256kFast$1 / $2$0.2Coding-oriented and the cheapest input rate here, but the smallest window — transcripts hit the wall sooner.
Bot brain

500k context · $2/$6 per 1M in/out · $0.5 cached. Above 200.0k tokens every token in the request re-bills at $4/$12.

Full append
$1.58
730.2k prompt tokens
Windowed
$0.7324
308.6k prompt tokens
Saved
54%
$1575.60 → $732.42 / 1k convs
Prompt tokens per turn append windowed
011.7k23.4kturn 1turn 60

Appending every turn makes the prompt grow linearly, so the total tokens billed across a conversation grow with the square of its length — turn 60 pays for turns 1-59 all over again. A verbatim window flattens that to a constant per turn. Push turns and reply length up until the amber line appears to see the 200k long-context cliff.

Before you launch

Guardrails

Per-conversation budget

Sum cost_usd for the conversation; over the ceiling, route to handoff_to_human.

One looping user can outspend a hundred normal ones.

Per-user message rate

Token bucket keyed on chat_id in Redis / Upstash.

Bots get spammed, and every spam message is a paid inference.

Step cap

stopWhen: isStepCount(5)

Every additional step resends the whole prompt at full price.

History cap

VERBATIM_TURNS + a rolling summary.

Prompt cost grows with turns otherwise — see the calculator above.

Output cap

maxOutputTokens sized to the surface.

Output tokens cost roughly 3x input; WhatsApp bubbles are small anyway.

Global kill switch

A feature flag checked before every generate call.

The only thing that actually stops a runaway at 2am.

codeAmani AI routing

Grok → Claude Fallback

Grok 4.6
primary
429 / 5xx
backoff ×4
Claude Sonnet 5
same tools
provider.ts
// lib/bot/provider.ts - one interface, two providers
import { xai } from "@ai-sdk/xai";
import { anthropic } from "@ai-sdk/anthropic";
import { generateText, stepCountIs } from "ai";

const PRIMARY = xai("grok-4.6");
const FALLBACK = anthropic("claude-sonnet-5");

export async function generateReply(args: Omit<Parameters<typeof generateText>[0], "model">) {
  try {
    return await withBackoff(() => generateText({ model: PRIMARY, ...args }));
  } catch (err) {
    // Same tools, same schema, same persona - only the brain changes
    console.warn("grok failed, falling back to claude", err);
    return generateText({ model: FALLBACK, ...args });
  }
}

The fallback only works if it is wired behind the same interface — same tool schemas, same persona, same message shape. A fallback you have never exercised is not a fallback; fire it deliberately in staging.

Official links

Resources

This portal covers the bot — the loop, the state, the money. For the Grok API surface itself (the full model catalogue, Live Search, image generation) see the xAI guide; for the messaging channel see WhatsApp Business API.

Prices are USD per 1M tokens, paid tier, verified Sep 1, 2026 against docs.x.ai/docs/models. Long-context rates apply to the whole request once the prompt reaches 200.0k tokens.