Anthropic Integration Guide
Technology: anthropic · Category: ai · Last reviewed: 2026-08-30
Source: https://tech-stack.codeamanilabs.org/guide/anthropic
Insight:
Claude is codeAmani's primary model for reasoning and code generation.
claude-opus-5($5/$25 per MTok) is the default — do not downgrade for cost without a measured reason. Step down toclaude-sonnet-5($2/$10) for everyday volume andclaude-haiku-4-5($1/$5) for latency-sensitive bulk work; step up toclaude-fable-5($10/$50) for frontier long-running agents. Tune spend withoutput_config.effortplus adaptive thinking, not the removedbudget_tokens. Call Claude through the Vercel AI SDK / AI Gateway (anthropic/...model strings) so provider swaps stay config, not code, and lean on prompt caching whenever a large system prompt or RAG context repeats across calls.
█████╗ ███╗ ██╗████████╗██╗ ██╗██████╗ ██████╗ ██████╗ ██╗ ██████╗
██╔══██╗████╗ ██║╚══██╔══╝██║ ██║██╔══██╗██╔═══██╗██╔══██╗██║██╔════╝
███████║██╔██╗ ██║ ██║ ███████║██████╔╝██║ ██║██████╔╝██║██║
██╔══██║██║╚██╗██║ ██║ ██╔══██║██╔══██╗██║ ██║██╔═══╝ ██║██║
██║ ██║██║ ╚████║ ██║ ██║ ██║██║ ██║╚██████╔╝██║ ██║╚██████╗
╚═╝ ╚═╝╚═╝ ╚═══╝ ╚═╝ ╚═╝ ╚═╝╚═╝ ╚═╝ ╚═════╝ ╚═╝ ╚═╝ ╚═════╝
Anthropic Integration Guide
Focus: Using Anthropic's APIs, SDKs, and MCP tooling directly within Claude Code workflows and automation pipelines.
Overview
Anthropic is the company behind Claude and Claude Code itself. Integrating Anthropic's APIs into Claude Code lets you build AI-assisted workflows, automate code generation, chain Claude API calls inside hooks, and extend Claude Code with custom MCP servers — all using first-party tooling.
Here is the big picture of how these first-party pieces fit together — once you see the shape, everything below slots right in.
flowchart TD
A["You · prompt or hook event"] --> B["Claude Code CLI"]
B --> C["Anthropic SDK · messages.create"]
C --> D["Claude API"]
D --> E["Response · code, review, tests"]
B --> F["MCP servers · custom tools"]
F --> B
B --> G["Automation · hooks and slash commands"]
G --> C
Official Documentation
| Resource | URL |
|---|---|
| Claude API Docs | https://platform.claude.com/docs/en/home |
| Claude Code Docs | https://code.claude.com/docs |
| Models overview | https://platform.claude.com/docs/en/models/overview |
| Managed Agents | https://platform.claude.com/docs/en/managed-agents/quickstart |
| CLI, SDKs & libraries | https://platform.claude.com/docs/en/cli-sdks-libraries/overview |
| Release notes | https://platform.claude.com/docs/en/release-notes/overview |
| Help Center (accounts, billing, plans) | https://support.claude.com/en |
| Model Context Protocol | https://modelcontextprotocol.io |
| API Reference | https://platform.claude.com/docs/en/api/overview |
| MCP SDK (TypeScript) | https://github.com/modelcontextprotocol/typescript-sdk |
| MCP SDK (Python) | https://github.com/modelcontextprotocol/python-sdk |
Docs domain moved. Anthropic's developer docs now live at
platform.claude.com/docs(the olddocs.anthropic.com/...URLs 301-redirect there); the console is atplatform.claude.com, status atstatus.claude.com, and pricing atclaude.com/pricing. Claude Code docs stay atcode.claude.com/docs.
MCP Server Setup
Claude Code as an MCP Server
Claude Code itself can act as an MCP server, exposing its tools to other clients.
# Start Claude Code as an MCP server (stdio transport)
claude mcp serve
Building a Custom MCP Server with @anthropic-ai/mcpb
@anthropic-ai/mcpb is Anthropic's official MCP bundle tool for creating distributable local MCP servers.
npm install -g @anthropic-ai/mcpb
Create a new MCP bundle project:
mcpb init my-server
cd my-server
mcpb build
mcpb install # installs the bundle into Claude Code
Connecting to the Official Claude Code MCP Server
# Add Claude Code as an MCP server inside another MCP client
claude mcp add claude-code -- claude mcp serve
.mcp.json Configuration
Create .mcp.json in your project root to auto-connect MCP servers when Claude Code opens:
{
"mcpServers": {
"anthropic-code": {
"command": "claude",
"args": ["mcp", "serve"],
"env": {}
}
}
}
Claude Code CLI Integration
Installation
npm install -g @anthropic-ai/claude-code
Key Commands
# Start interactive session
claude
# Run a one-shot prompt (non-interactive)
claude -p "Explain the auth flow in src/auth.ts"
# Run with a specific model
claude --model claude-opus-5
# Continue the most recent session
claude --continue
# Run a bash command within a Claude session
claude -p "Fix the TypeScript errors" --allowedTools Bash,Edit,Write
# Start as MCP server
claude mcp serve
# Manage MCP servers
claude mcp add <name> -- <command> [args]
claude mcp list
claude mcp remove <name>
# Add remote MCP server (HTTP transport)
claude mcp add --transport http my-server https://my-server.example.com/mcp
Anthropic SDK Integration
A single messages.create call is the heartbeat of every SDK example below — here is exactly what happens on each request.
sequenceDiagram
participant App as "Your app or script"
participant SDK as "Anthropic SDK"
participant API as "Claude API"
App->>SDK: "messages.create · model, max_tokens, messages"
SDK->>API: "authenticated request · ANTHROPIC_API_KEY"
API-->>SDK: "message · content blocks"
SDK-->>App: "message.content·0·.text"
Node.js / TypeScript
npm install @anthropic-ai/sdk
import Anthropic from "@anthropic-ai/sdk";
// Zero-arg constructor resolves ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN,
// or an `ant auth login` profile — see Authentication below.
const client = new Anthropic();
const message = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
messages: [{ role: "user", content: "Review this code for security issues." }],
});
// `content` is a discriminated union — narrow by `.type` before reading `.text`.
for (const block of message.content) {
if (block.type === "text") console.log(block.text);
}
The TypeScript SDK is at
@anthropic-ai/sdk0.122.0. Themessages.createshape above is stable across the 0.x line.
Python
pip install anthropic
import anthropic
# Resolves ANTHROPIC_API_KEY, ANTHROPIC_AUTH_TOKEN, or an `ant auth login` profile.
client = anthropic.Anthropic()
message = client.messages.create(
model="claude-opus-5",
max_tokens=16000,
messages=[{"role": "user", "content": "Generate unit tests for this function."}],
)
# Narrow by block type — tool_use blocks have no .text attribute.
for block in message.content:
if block.type == "text":
print(block.text)
The Python SDK is at
anthropic1.2.0. The 0.x → 1.x breaking change was the upgrade to httpx 2 (see the SDK'sMIGRATION.md);client.messages.create(...)usage is unchanged. Pinanthropic>=1,<2and make sure your environment allows httpx 2.
Thinking, Effort, and the Claude 5 Request Shape
The Claude 5 family changed the request surface in ways that silently break code written for the 4.x line. Three things are now rejected with a 400 on claude-fable-5, claude-opus-5, and claude-sonnet-5:
| Removed | Replacement |
|---|---|
thinking: { type: "enabled", budget_tokens: N } |
thinking: { type: "adaptive" } |
temperature / top_p / top_k |
none — sampling is model-managed |
| Assistant-message prefill (pre-filling the last turn) | output_config.format, or a system instruction |
Adaptive thinking lets Claude decide when and how deeply to reason instead of you pre-buying a token budget. On claude-opus-5 thinking is on by default — omitting the parameter runs adaptive. Effort is the dial that replaced budget_tokens, and it lives inside output_config, not at the top level:
const message = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
thinking: { type: "adaptive", display: "summarized" },
output_config: { effort: "xhigh" }, // low | medium | high | xhigh | max
messages: [{ role: "user", content: "Refactor this module and explain why." }],
});
effortdefaults tohigh.xhighis the sweet spot for coding and agentic work;lowsuits subagents and simple classification.displaydefaults to"omitted"on Claude 5 models — thinking blocks stream with empty text. If you surface reasoning in a UI, setdisplay: "summarized"explicitly, or users see a long silent pause before any output.- Thinking is billed identically under every
displaysetting; the raw chain of thought is never returned.
max_tokensis a truncation cliff, not a cost control. Default to16000for non-streaming calls and64000when streaming. Claude 5 models support up to 128K output tokens, but the SDKs require streaming at that size to avoid HTTP timeouts. The old1024habit truncates mid-thought and buys you a retry.
Structured Outputs
When you need JSON that actually validates, constrain the response instead of parsing hopefully. Use output_config.format — the older top-level output_format parameter is deprecated:
const message = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
output_config: {
format: {
type: "json_schema",
schema: {
type: "object",
properties: {
severity: { type: "string", enum: ["low", "medium", "high"] },
summary: { type: "string" },
},
required: ["severity", "summary"],
additionalProperties: false,
},
},
},
messages: [{ role: "user", content: "Triage this Sentry error: " + stackTrace }],
});
For tool arguments, set strict: true as a top-level field on the tool definition (not on tool_choice); the schema needs additionalProperties: false plus required. Structured outputs are incompatible with document citations — sending both returns a 400.
Managed Agents
The docs home now leads with two developer surfaces: the Messages API (you own the loop) and Managed Agents (Anthropic runs the loop and hosts the sandbox where tools execute). Reach for Managed Agents when the alternative is writing your own scheduler, session store, and container runtime.
The flow is agent once → session per run. model, system, and tools live on the agent, never on the session:
// 1. Create the agent once. Store the ID — never call this in the request path.
const agent = await client.beta.agents.create({
model: "claude-opus-5",
system: "You reconcile M-Pesa settlement files against Stripe payouts.",
tools: [{ type: "bash_20250124", name: "bash" }],
});
// 2. Start a session per run, referencing the stored agent ID.
const session = await client.beta.sessions.create({ agent_id: agent.id });
- The beta header
managed-agents-2026-04-01is set automatically by the SDK forclient.beta.{agents,sessions,environments,vaults,deployments}.*. - Scheduled deployments fire sessions on a cron cadence — use those rather than a client-side scheduler for nightly or weekly agent jobs.
- Vault credentials (
environment_variable) are held by Anthropic and substituted at egress, so secrets never enter the sandbox. Prefer them to passing keys into tool code. - Not available on Bedrock, Vertex AI, or Foundry — use Messages + tool use on those platforms.
Three different things that sound alike. Agent Skills generate
.pptx/.xlsxviacontainer.skillson a normalmessages.create. The Claude Agent SDK (@anthropic-ai/claude-agent-sdk) is Claude Code packaged as a library that you host. Only Managed Agents supplies both a managed harness and managed deployment.
Prompt Caching
When you reuse the same large block across calls — a frozen system prompt, a long tool set, retrieved RAG context — mark it with cache_control: { type: "ephemeral" }. Anthropic caches that prefix and serves it back at roughly 0.1× input cost on cache hits, with lower latency. For codeAmani's AI features (review bots, support agents, Swahili/English assistants) this is the single biggest cost lever when the per-request question is small but the shared context is huge.
The one rule: caching is a prefix match. Render order is tools → system → messages, and any byte change before a breakpoint invalidates everything after it. Keep stable content first; put volatile content (the user's question, a timestamp, a per-request ID) after the last breakpoint.
Two ways to cache. Automatic caching — set one
cache_control: { type: "ephemeral" }at the top level of the request and Anthropic manages the breakpoint, moving it forward to the last cacheable block as a conversation grows (ideal for multi-turn chat/agents). Explicit breakpoints — the block-level markers shown below, for fine-grained control over exactly what caches. Automatic caching consumes one of your 4 breakpoint slots. The explicit form is used in this example because codeAmani's hot paths reuse a fixed tool set + system prompt.
import Anthropic from "@anthropic-ai/sdk";
const client = new Anthropic();
const message = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
// Cache the tool set — tools render at position 0, so this prefix is reused first.
tools: [
{
name: "search_orders",
description: "Look up M-Pesa orders by phone number.",
input_schema: {
type: "object",
properties: { phone: { type: "string" } },
required: ["phone"],
},
cache_control: { type: "ephemeral" },
},
],
// Cache the large, frozen system prompt — breakpoint on the LAST block caches tools + system together.
system: [
{
type: "text",
text: LARGE_SHARED_PROMPT, // e.g. product catalog, brand rules, RAG context
cache_control: { type: "ephemeral" }, // add `ttl: "1h"` for bursty traffic with idle gaps
},
],
// Volatile content goes last, after the cached prefix — no marker here.
messages: [{ role: "user", content: "Where is my last order?" }],
});
// Confirm it worked — cache_read_input_tokens should be > 0 on the 2nd+ identical-prefix call.
console.log(message.usage.cache_read_input_tokens, message.usage.cache_creation_input_tokens);
Notes that bite in practice:
- Max 4 breakpoints per request. Minimum cacheable prefix depends on model — 512 tokens (Opus 5, Fable 5), 1024 (Sonnet 5), 4096 (Haiku 4.5). Shorter prefixes silently won't cache (
cache_creation_input_tokens: 0, no error). - Verify with
usage. Ifcache_read_input_tokensstays 0 across repeated calls, a silent invalidator is in the prefix —new Date()/Date.now()in the system prompt, unsortedJSON.stringify, a per-user ID interpolated early, or a tool set that changes per request. - Don't interpolate dynamic values into the system prompt (current date, user name, mode). Those sit at the front and break every downstream cache — pass them in a later
messagesentry instead. - Economics: writes cost ~1.25× (5m TTL) or ~2× (1h TTL). Break-even is two calls for the 5-minute default, ~three for the 1-hour TTL.
Via the AI Gateway: when calling Claude through the Vercel AI SDK with
anthropic/...model strings (codeAmani's default — see the insight), passcache_controlthroughproviderOptions.anthropicso the marker reaches the underlying API. Caching is an Anthropic-side feature; the gateway forwards it but does not invent it.
flowchart LR
A["Request · tools then system then messages"] --> B{"Prefix byte-identical<br/>to a cached entry"}
B -->|"yes · cache hit"| C["Served at 0.1× input cost<br/>cache_read_input_tokens > 0"]
B -->|"no · cache miss"| D["Full price · writes cache<br/>at 1.25× then reusable"]
D --> E["Next call reuses the prefix"]
E --> B
Canonical reference: https://platform.claude.com/docs/en/build-with-claude/prompt-caching
Authentication and Environment Variables
An unset ANTHROPIC_API_KEY does not mean you have no credentials. The SDKs and the ant CLI resolve in this order, first match wins:
ANTHROPIC_API_KEYANTHROPIC_AUTH_TOKEN- The
ANTHROPIC_PROFILE-selected (or active) OAuth profile fromant auth login - Workload Identity Federation environment variables
- The default profile on disk
A bare new Anthropic() therefore works after ant auth login with no env var set. Check which source is actually live before concluding a key is missing:
ant auth status # shows the active credential source and profile
ant auth login # stores a profile under ~/.config/anthropic/ that the SDKs read
# Server-side only — never ship this into a client bundle
ANTHROPIC_API_KEY=sk-ant-...
# Optional overrides
ANTHROPIC_BASE_URL=https://api.anthropic.com # default
ANTHROPIC_MODEL=claude-opus-5 # default model for the claude CLI
ANTHROPIC_PROFILE=work # select a named `ant auth login` profile
# Claude Code specific
CLAUDE_CODE_MAX_OUTPUT_TOKENS=32000
Set these in your shell profile, in .env.local for local dev, or in the Vercel dashboard for production.
Raw
curlunder OAuth: anantprofile is not an API key. Mint a short-lived token withant auth print-credentials --access-token, then send it asAuthorization: Bearer <token>plus the headeranthropic-beta: oauth-2025-04-20. OAuth tokens do not go inx-api-key— converting a working curl from an API key is a header change, not just a value swap.
Automation Workflows
Claude Code Hooks
Hooks run shell commands automatically at lifecycle events. Configure in .claude/settings.json:
{
"hooks": {
"PreToolUse": [
{
"matcher": "Bash",
"hooks": [
{
"type": "command",
"command": "echo 'Tool: Bash about to run' >> .claude/audit.log"
}
]
}
],
"PostToolUse": [
{
"matcher": "Write",
"hooks": [
{
"type": "command",
"command": "npx prettier --write $CLAUDE_FILE_PATH 2>/dev/null || true"
}
]
}
],
"Stop": [
{
"hooks": [
{
"type": "command",
"command": "node scripts/notify-complete.js"
}
]
}
]
}
}
Slash Commands
Create custom slash commands as markdown files in .claude/commands/:
mkdir -p .claude/commands
.claude/commands/review.md:
Review the following code for: security issues, performance problems, and code quality.
Focus on: $ARGUMENTS
Provide actionable fixes with code examples.
Usage inside Claude Code: /project:review src/api/auth.ts
Headless Automation with the SDK
Use the Claude API to automate code review in CI:
// scripts/ai-review.ts
import Anthropic from "@anthropic-ai/sdk";
import { readFileSync } from "fs";
const client = new Anthropic();
const diff = readFileSync("latest.diff", "utf-8");
const review = await client.messages.create({
model: "claude-opus-5",
max_tokens: 16000,
output_config: { effort: "xhigh" },
system: "You are a senior code reviewer. Be concise and actionable.",
messages: [{ role: "user", content: `Review this diff:\n\n${diff}` }],
});
for (const block of review.content) {
if (block.type === "text") console.log(block.text);
}
Common Use Cases
| Use Case | Approach |
|---|---|
| Automated PR review | Fetch diff via gh, pipe to Claude API |
| Code generation | claude -p "Generate CRUD endpoints for User model" |
| Test generation | Hook on PostToolUse[Write] to auto-generate tests |
| Documentation | claude -p "Document all exported functions in src/" |
| Security scanning | Combine with Semgrep output piped to Claude API |
| Refactoring | Use --continue sessions for multi-step refactors |
CLAUDE.md Configuration
Create CLAUDE.md at your project root to give Claude Code persistent context:
# Project Context
## Tech Stack
- TypeScript, Node.js 22, PostgreSQL
- Test runner: Vitest
- Linter: ESLint + Prettier
## Conventions
- Use `async/await` — no raw Promises
- All functions must have JSDoc comments
- Tests go in `__tests__/` next to source files
## Forbidden
- Never use `any` type
- Never commit `.env` files
Troubleshooting
| Issue | Fix |
|---|---|
ANTHROPIC_API_KEY not found |
Export it in shell: export ANTHROPIC_API_KEY=sk-ant-... |
| Rate limit errors | Add retry logic with exponential backoff |
| MCP server not connecting | Run claude mcp list to verify registration |
| Hooks not firing | Check .claude/settings.json syntax with cat .claude/settings.json | jq . |
| Model not available | Check the live list at platform.claude.com/docs/en/models/overview |
400 on budget_tokens |
Removed on Claude 5 — use thinking: { type: "adaptive" } + output_config.effort |
400 on temperature or a prefilled assistant turn |
Both removed on Claude 5 — shape output with output_config.format |
| Reasoning UI shows a long blank pause | display defaults to "omitted" — set thinking.display: "summarized" |
Official docs: