Grok Bot
Developer Portal
A bot is not an API call — it is a stateful loop wrapped around a stateless model. The model is one of six jobs; these are the other five, and what each one costs.
The Bot Loop
Recompute the HMAC over the RAW request body and compare in constant time. Read req.text() first — parsing to JSON before verifying changes the bytes and the signature will never match.
Skip it and anyone who learns your webhook URL can post messages as any user. This is the only thing standing between your bot and an open relay.
// app/api/whatsapp/webhook/route.ts
import { after } from "next/server";
import { handleTurn } from "@/lib/bot/turn";
export async function POST(req: Request) {
// Signature is computed over the RAW body - read text, never req.json() first
const raw = await req.text();
if (!verifySignature(raw, req.headers.get("x-hub-signature-256"))) {
return new Response("invalid signature", { status: 401 });
}
const payload = JSON.parse(raw);
const msg = payload.entry?.[0]?.changes?.[0]?.value?.messages?.[0];
// ACK first, work after - Meta retries anything it does not get a fast 200 on
if (msg) after(() => handleTurn(msg));
return new Response("ok", { status: 200 });
}Chat Surfaces
WhatsApp Cloud API
GA- Transport
- REST over Graph API (no SDK)
- Streaming
- None — one message per send
- Message cap
- ~4096 chars
- Webhook
- Fast 200 or Meta retries
Outside 24h from the user's last message, only pre-approved template messages are delivered. Free-form replies are rejected.
Telegram
- Transport
- Bot API via grammy
- Streaming
- Simulated (editMessageText)
- Message cap
- 4096 chars
- Webhook
- Standard webhook
Telegram tells you exactly how long to wait in retry_after — honour it rather than applying your own backoff.
Discord
- Transport
- Gateway + REST via discord.js
- Streaming
- Simulated (message edits)
- Message cap
- 2000 chars
- Webhook
- 3-second interaction deadline
You must acknowledge an interaction within 3 seconds. Any model call is slower than that, so deferReply() is mandatory, not optional.
Getting Started
npm install ai @ai-sdk/xai @ai-sdk/anthropic zod
npm install grammy # Telegram
# WhatsApp Cloud API is plain REST - no SDK neededEvery one of these is server-side only. An XAI_API_KEY in a NEXT_PUBLIC_* variable is shipped to every browser that loads the page.
Conversation State
Keep the last N turns verbatim and compress everything older into one paragraph the persona can read. The window is the single highest-leverage cost control you have — the calculator below prices it.
// lib/bot/history.ts
import type { ModelMessage } from "ai";
const VERBATIM_TURNS = 12;
export async function buildMessages(conversationId: string): Promise<ModelMessage[]> {
const convo = await getConversation(conversationId);
const recent = await getRecentMessages(conversationId, VERBATIM_TURNS);
const messages: ModelMessage[] = [];
if (convo.summary) {
messages.push({
role: "user",
content: `[Earlier in this conversation]\n${convo.summary}`,
});
}
for (const m of recent) {
messages.push({ role: m.role, content: m.content } as ModelMessage);
}
return messages;
}Persona Design
Scope fence
What the bot is for, and explicitly what it will not do. First, because it is the instruction most likely to survive a long transcript.
Voice
Tone, length, language. For codeAmani bots this is where Swahili / Sheng register and code-switching rules live.
Hard rules
Never invent prices. Never promise delivery dates. Never repeat a customer's full phone number back to them.
Escalation
The exact conditions for calling handoff_to_human — anger, refunds, anything involving money moving.
Version the persona like code. A prompt change is a behaviour change, and you will want to know which release started apologising in the wrong language.
Tool Loop
Scope every tool to the caller
Take chat_id from experimental_context, never from a model-supplied argument. The model will happily pass someone else's id.
Validate with a schema, not trust
inputSchema is a boundary. Treat tool arguments exactly like untrusted form input, because that is what they are.
Cap the steps
stopWhen: stepCountIs(5). Each step is a full prompt resend at full price.
// lib/bot/tools.ts
import { tool } from "ai";
import { z } from "zod";
export const botTools = {
checkOrder: tool({
description: "Look up an order by its reference for the CURRENT customer.",
inputSchema: z.object({ reference: z.string().min(4) }),
// The model asks. Your code decides - and scopes to the caller.
execute: async ({ reference }, { experimental_context }) => {
const { chatId } = experimental_context as { chatId: string };
const order = await db.order.findFirst({ where: { reference, chatId } });
return order ?? { error: "not_found" };
},
}),
handoffToHuman: tool({
description: "Escalate to a human agent when the customer asks or is upset.",
inputSchema: z.object({ reason: z.string() }),
execute: async ({ reason }, { experimental_context }) => {
const { chatId } = experimental_context as { chatId: string };
await notifyAgents({ chatId, reason });
return { escalated: true };
},
}),
};Cost per Conversation
Tokens per request is the wrong unit for a bot. The unit is cost per conversation, because that is what scales with users. Drag the sliders to price your own bot.
| Model | Context | Type | In / Out | Cached in | Why you would pick it |
|---|---|---|---|---|---|
Grok 4.6GA grok-4.6 | 500k | Fast | $2 / $6 | $0.5 | xAI's current general-purpose recommendation and the sensible default bot brain. |
Grok 4.5GA grok-4.5 | 500k | Fast | $2 / $6 | $0.3 | Previous generation. Same headline price as 4.6 with a cheaper cached-input rate. |
Grok 4.3GA grok-4.3 | 1M | Fast | $1.25 / $2.5 | $0.2 | Cheapest per token with the largest context. The value pick for high-volume bots. |
Grok 4.20 ReasoningGA grok-4.20-0309-reasoning | 1M | Reasoning | $1.25 / $2.5 | $0.2 | Reasoning variant. Same token price as 4.3 but latency is the tax — often too slow for a chat surface. |
Grok Build 0.1GA grok-build-0.1 | 256k | Fast | $1 / $2 | $0.2 | Coding-oriented and the cheapest input rate here, but the smallest window — transcripts hit the wall sooner. |
500k context · $2/$6 per 1M in/out · $0.5 cached. Above 200.0k tokens every token in the request re-bills at $4/$12.
Appending every turn makes the prompt grow linearly, so the total tokens billed across a conversation grow with the square of its length — turn 60 pays for turns 1-59 all over again. A verbatim window flattens that to a constant per turn. Push turns and reply length up until the amber line appears to see the 200k long-context cliff.
Guardrails
Per-conversation budget
Sum cost_usd for the conversation; over the ceiling, route to handoff_to_human.One looping user can outspend a hundred normal ones.
Per-user message rate
Token bucket keyed on chat_id in Redis / Upstash.Bots get spammed, and every spam message is a paid inference.
Step cap
stopWhen: isStepCount(5)Every additional step resends the whole prompt at full price.
History cap
VERBATIM_TURNS + a rolling summary.Prompt cost grows with turns otherwise — see the calculator above.
Output cap
maxOutputTokens sized to the surface.Output tokens cost roughly 3x input; WhatsApp bubbles are small anyway.
Global kill switch
A feature flag checked before every generate call.The only thing that actually stops a runaway at 2am.
Grok → Claude Fallback
// lib/bot/provider.ts - one interface, two providers
import { xai } from "@ai-sdk/xai";
import { anthropic } from "@ai-sdk/anthropic";
import { generateText, stepCountIs } from "ai";
const PRIMARY = xai("grok-4.6");
const FALLBACK = anthropic("claude-sonnet-5");
export async function generateReply(args: Omit<Parameters<typeof generateText>[0], "model">) {
try {
return await withBackoff(() => generateText({ model: PRIMARY, ...args }));
} catch (err) {
// Same tools, same schema, same persona - only the brain changes
console.warn("grok failed, falling back to claude", err);
return generateText({ model: FALLBACK, ...args });
}
}The fallback only works if it is wired behind the same interface — same tool schemas, same persona, same message shape. A fallback you have never exercised is not a fallback; fire it deliberately in staging.
Resources
This portal covers the bot — the loop, the state, the money. For the Grok API surface itself (the full model catalogue, Live Search, image generation) see the xAI guide; for the messaging channel see WhatsApp Business API.
Prices are USD per 1M tokens, paid tier, verified Sep 1, 2026 against docs.x.ai/docs/models. Long-context rates apply to the whole request once the prompt reaches 200.0k tokens.