Google AI Studio (Gemini API) Integration Guide
Gemini API Developer Portal
Every model, endpoint, and price — audited Sep 1, 2026
What is the Gemini API?
Gemini 3 Pro/Flash, 1M-token context, multimodal-native, with code execution.
The 1M-token context is the structural unlock: an entire codebase, a 90-page PDF, or hours of audio in one call. Gemini also has built-in `code_execution` tool (the model writes and runs Python, sees the result, iterates). Free tier on AI Studio is generous for prototypes — perfect for evaluating against Claude/GPT before committing. For codeAmani, Gemini 3 Flash is the cheap multimodal workhorse (receipts, ID extraction, transcription) while Claude handles complex reasoning.
Five Gemini primitives
Same API for text, image, audio and video — and a 1M-token context to fit them all.
██████╗ ██████╗ ██████╗ ██████╗ ██╗ ███████╗ █████╗ ██╗
██╔════╝ ██╔═══██╗██╔═══██╗██╔════╝ ██║ ██╔════╝ ██╔══██╗██║
██║ ███╗██║ ██║██║ ██║██║ ███╗██║ █████╗ ███████║██║
██║ ██║██║ ██║██║ ██║██║ ██║██║ ██╔══╝ ██╔══██║██║
╚██████╔╝╚██████╔╝╚██████╔╝╚██████╔╝███████╗███████╗ ██║ ██║██║
╚═════╝ ╚═════╝ ╚═════╝ ╚═════╝ ╚══════╝╚══════╝ ╚═╝ ╚═╝╚═╝
███████╗████████╗██╗ ██╗██████╗ ██╗ ██████╗
██╔════╝╚══██╔══╝██║ ██║██╔══██╗██║██╔═══██╗
███████╗ ██║ ██║ ██║██║ ██║██║██║ ██║
╚════██║ ██║ ██║ ██║██║ ██║██║██║ ██║
███████║ ██║ ╚██████╔╝██████╔╝██║╚██████╔╝
╚══════╝ ╚═╝ ╚═════╝ ╚═════╝ ╚═╝ ╚═════╝Google AI Studio (Gemini API) Integration Guide
Focus: Google AI Studio is where codeAmani Labs creates its Gemini API key. The Gemini API powers text, multimodal, and image generation (it generates the dashboard's tech-stack thumbnails). This is the AI Studio / Developer-API path — distinct from Vertex AI.
Overview
Here's the high-level path your call takes — once you picture it, the rest of the guide clicks into place.
Google AI Studio issues a Gemini Developer API key
that authenticates calls to generativelanguage.googleapis.com. The official, current
SDK is @google/genai (JS/TS, 2.20.0 as of 2026-09-01, needs Node 20+) and
google-genai (Python, 2.21.0, needs Python 3.10+). The older
@google/generative-ai package is legacy/deprecated (last release 0.24.1,
Apr 2025) and no longer receives new Gemini features — do not use it for new code.
Gemini 3 is the current model generation (as of 2026-09-01); the Gemini 2.5 line
is the previous generation and still available. Flash IDs iterate quickly
(gemini-3.5-flash → 3.6 → 3.7-flash), so pin a dated ID for production or use the
gemini-flash-latest alias to auto-track the newest release.
Official Documentation
| Resource | URL |
|---|---|
| Gemini API docs | https://ai.google.dev/gemini-api/docs |
| Gemini 3 developer guide | https://ai.google.dev/gemini-api/docs/gemini-3 |
| Get an API key | https://ai.google.dev/gemini-api/docs/api-key |
| Google AI Studio | https://aistudio.google.com |
JS SDK (@google/genai) | https://googleapis.github.io/js-genai/ |
Python SDK (google-genai) | https://googleapis.github.io/python-genai/ |
| Image generation | https://ai.google.dev/gemini-api/docs/image-generation |
| Rate limits & tiers | https://ai.google.dev/gemini-api/docs/rate-limits |
| Error codes / troubleshooting | https://ai.google.dev/gemini-api/docs/troubleshooting |
1. Get an API key
- Go to Google AI Studio and sign in.
- Select Get API key → Create API key (in a Google Cloud project).
- Store it as
GEMINI_API_KEY(the SDK also readsGOOGLE_API_KEY).- codeAmani convention:
.env.local(gitignored) for local dev + the Vercel project's env vars for deploys. Server-side only — never ship the key to the browser.
- codeAmani convention:
A Maps Platform API key is NOT a Gemini key: calling the Gemini API with one returns
403 API_KEY_SERVICE_BLOCKED. Use a key created in AI Studio.
2. Install the SDK
npm install @google/genai # JS/TS
pip install google-genai # Python3. Generate text
import { GoogleGenAI } from "@google/genai";
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const res = await ai.models.generateContent({
model: "gemini-3.7-flash",
contents: "Explain M-Pesa STK Push in one sentence.",
});
console.log(res.text);4. Generate images
Image generation now runs entirely through the Gemini "Nano Banana" image models
via generateContent — the dedicated Imagen models (imagen-4.0-generate-001
/ -ultra- / -fast-) and the old ai.models.generateImages path were shut down on
2026-08-17. Pick the tier by quality vs cost.
All three tiers call generateContent and return the image as an inline-data part —
only the model ID changes. Per-image output prices (paid tier, verified 2026-09-01):
| Tier | Model ID | Price / image |
|---|---|---|
| Nano Banana Pro — world knowledge, brand consistency, up to 4K | gemini-3-pro-image | $0.134 (1K/2K) · $0.24 (4K) |
| Nano Banana 2 — balanced generalist workhorse | gemini-3.1-flash-image | $0.067 (1K) · $0.101 (2K) · $0.151 (4K) |
| Nano Banana 2 Lite — cheapest, ultra-low latency (GA 2026-06-30) | gemini-3.1-flash-lite-image | $0.0336 (1K) |
const res = await ai.models.generateContent({
model: "gemini-3.1-flash-image", // or "gemini-3-pro-image" for premium
contents: "A glossy 3D emblem of a green database with a lightning bolt",
});
const part = res.candidates?.[0]?.content?.parts?.find((p) => p.inlineData);
const bytes = part?.inlineData?.data; // base64 PNGREST — the same call this repo's scripts/generate-thumbnails.py uses to build the
card thumbnails (it currently pins gemini-2.5-flash-image, the original Nano Banana —
still available, now the legacy tier):
curl "https://generativelanguage.googleapis.com/v1beta/models/gemini-3.1-flash-image:generateContent" \
-H "x-goog-api-key: $GEMINI_API_KEY" -H "Content-Type: application/json" \
-d '{"contents":[{"parts":[{"text":"A colorful app-icon logo"}]}]}'
# response: candidates[].content.parts[].inlineData.data (base64 PNG)Models (verified 2026-09-01)
| Model | Use |
|---|---|
gemini-3.7-flash | current latest stable flash — text / multimodal reasoning, 1M-token context |
gemini-3.6-flash | previous stable flash — same promo pricing as 3.7 |
gemini-3.5-flash | legacy flash, still GA — routine high-throughput work |
gemini-3.5-flash-lite | cheapest 3.5-line model |
gemini-3.1-flash-lite | lowest-cost / highest-QPS text model |
gemini-3.1-pro-preview | Gemini 3 Pro — highest-capability reasoning (preview) |
gemini-3-pro-image | premium image gen/edit — "Nano Banana Pro" (up to 4K) |
gemini-3.1-flash-image | workhorse image gen/edit — "Nano Banana 2" |
gemini-3.1-flash-lite-image | cheapest image gen — "Nano Banana 2 Lite" (GA 2026-06-30) |
gemini-2.5-flash-image | legacy image model — original "Nano Banana" (still available) |
gemini-3.5-transcribe | speech-to-text with diarization + language detection (stable) |
gemini-2.5-flash / gemini-2.5-pro | previous-generation text models (still available) |
Retired: all Imagen 4 IDs (
imagen-4.0-generate-001/-ultra-/-fast-) and Imagen 3 were shut down 2026-08-17 — migrate to the Nano Banana models above. Flash IDs iterate fast; re-verify from the models page or usegemini-flash-latest.
List live models for a key:
GET https://generativelanguage.googleapis.com/v1beta/models with header
x-goog-api-key: $GEMINI_API_KEY.
Errors, rate limits & retries
The Gemini API returns standard HTTP codes with a canonical status name. The ones worth retrying are transient (rate limit + server-side); the rest are bugs in your request and retrying just wastes quota.
| Code | Status | Meaning | Retry? |
|---|---|---|---|
400 | INVALID_ARGUMENT | malformed request / bad field | No — fix the call |
403 | PERMISSION_DENIED | wrong/blocked key (e.g. a Maps key) | No — fix the key |
429 | RESOURCE_EXHAUSTED | you exceeded the rate limit / quota | Yes — backoff |
500 | INTERNAL | unexpected error on Google's side | Yes — backoff |
503 | UNAVAILABLE | service temporarily overloaded / down | Yes — backoff |
504 | DEADLINE_EXCEEDED | request didn't finish in time | Raise client timeout / shrink prompt |
Authoritative tables: the troubleshooting page (error codes) and the rate-limits page (tiers). The official docs do not prescribe a backoff algorithm, so the snippet below is a standard exponential-backoff-with-jitter pattern applied to the documented retryable codes.
The @google/genai SDK throws an ApiError that extends Error with a .status
field holding the HTTP code — so you branch on .status, not on string matching.
import { GoogleGenAI, ApiError } from "@google/genai";
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY });
const RETRYABLE = new Set([429, 500, 503]); // RESOURCE_EXHAUSTED, INTERNAL, UNAVAILABLE
const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms));
/** Run a Gemini call with exponential backoff + jitter on transient errors. */
async function withBackoff<T>(fn: () => Promise<T>, maxRetries = 5): Promise<T> {
for (let attempt = 0; ; attempt++) {
try {
return await fn();
} catch (err) {
const status = err instanceof ApiError ? err.status : undefined;
if (attempt >= maxRetries || status === undefined || !RETRYABLE.has(status)) {
throw err; // out of retries, or a non-retryable error like 400/403
}
// 1s, 2s, 4s, 8s ... capped at 30s, plus up to 1s of jitter
const delay = Math.min(2 ** attempt * 1000, 30_000) + Math.random() * 1000;
await sleep(delay);
}
}
}
const res = await withBackoff(() =>
ai.models.generateContent({
model: "gemini-3.7-flash",
contents: "Explain M-Pesa STK Push in one sentence.",
}),
);
console.log(res.text);Free tier vs paid. The free tier has tight per-minute and per-day quotas; once you enable billing your project moves to a paid usage tier with much higher limits. Exact RPM/TPD/RPD numbers vary by model and tier and change over time, so do not hard-code them — read your project's live limits in Google AI Studio and on the rate-limits page. For the thumbnail pipeline, image generation is metered separately and per-image, so a single 429 burst on the free tier is common — backoff plus caching in R2 keeps it cheap.
Gotcha: retrying a
400/403is pointless and, with a429, a tight retry loop with no backoff just digs the quota hole deeper — each rejected call can still count against your rate budget. Only retry the codes in the table above, always with growing delays, and cap total attempts.
codeAmani notes
- AI routing: Gemini is the image/multimodal provider here; Anthropic Claude remains
primary for reasoning/codegen (see
AI_WORKFLOWS.md). - Security: keep
GEMINI_API_KEYserver-side; call from API routes / scripts, never inline in client components. Restrict the key in Google Cloud where possible. - Cost: image generation is billed per image — generate thumbnails once and cache
them (this repo stores them in the
tech-stack-bucketR2 bucket).