AI Video Generation Integration Guide
Technology: ai-video · Category: ai · Last reviewed: 2026-08-23
Source: https://tech-stack.codeamanilabs.org/guide/ai-video
Insight:
Generative video is an async job, not a request/response: you submit a prompt, get a job id, then poll or receive a webhook before downloading a clip that costs cents-to-dollars per output-second and whose URL expires within days. The market moved fast through 2026 — Veo 3.1 is the quality leader, aggregators (fal
fal-ai/veo3.1, Replicategoogle/veo-3.1) let you swap models with one string, and OpenAI is exiting: the Sora 2 API shuts down 2026-09-24 with no announced successor. For codeAmani it is a distinct AI-routing modality — default to short 9:16 vertical clips for East African social / WhatsApp marketing, draft cheap on a fast tier, and re-host every render to R2 before the provider URL expires.
█████╗ ██╗ ██╗ ██╗██╗██████╗ ███████╗ ██████╗
██╔══██╗██║ ██║ ██║██║██╔══██╗██╔════╝██╔═══██╗
███████║██║ ██║ ██║██║██║ ██║█████╗ ██║ ██║
██╔══██║██║ ╚██╗ ██╔╝██║██║ ██║██╔══╝ ██║ ██║
██║ ██║██║ ╚████╔╝ ██║██████╔╝███████╗╚██████╔╝
╚═╝ ╚═╝╚═╝ ╚═══╝ ╚═╝╚═════╝ ╚══════╝ ╚═════╝
AI Video Generation Integration Guide
Focus: generative video is an async job, not a request/response. You submit a prompt, get a job id, then either poll the status or let a webhook call you back, and finally download the rendered clip. It is the exact reverse-API flow the repo already documents for webhooks — applied to a render that takes 30 seconds to several minutes.
Overview
Text-to-video and image-to-video models turn a prompt (and optionally a starting image) into a short MP4 with — increasingly — native audio. Unlike an LLM call that streams tokens back in seconds, a video render is long-running and expensive: a single 8-second 1080p clip can take a minute or two of GPU time and cost anywhere from a few cents to several dollars. Because of that, no serious provider makes you hold an HTTP connection open for the render. They all converged on the same shape:
- Submit the prompt → you immediately get a job / prediction / operation id.
- Wait for completion via one of two mechanisms:
- Poll — GET the job id every few seconds until
statusis terminal. - Webhook — register a callback URL; the provider POSTs you when the clip is ready (this is the cheaper, more scalable path, and it is the same receiver pattern documented in
webhooks/).
- Poll — GET the job id every few seconds until
- Download the rendered clip from the signed URL in the result (the file usually expires — re-host it to your own R2/Supabase bucket).
This guide treats AI video as a new modality in the codeAmani AI routing policy — alongside Claude (reasoning), OpenAI (structured), and HuggingFace (open models) — and shows how to call it in production through both first-party APIs (Google Veo via the Gemini API) and aggregators (Replicate, fal.ai) that put dozens of models behind one async interface.
Models & providers
Prices are per second of output and move fast — always re-check the model page before committing a budget. As of the review date (2026-08-23):
| Model | Best at | Native audio | Duration / res | ~Cost/sec | How to call |
|---|---|---|---|---|---|
| Google Veo 3.1 | Top-tier realism, lip-sync, 1080p/4K | Yes (always on) | 4/6/8s · 720p/1080p/4K | ~$0.40 (Standard); ~$0.15 Fast | Gemini API (@google/genai), Vertex AI, Replicate (google/veo-3.1), fal (fal-ai/veo3.1) |
| Runway Gen-4.5 | Cinematic control, native 4K, prompt adherence | No (Aleph 2.0 edits w/ audio) | short clips · up to 4K | ~$0.20 | Runway Dev API (@runwayml/sdk), aggregators |
| Luma Ray3 / Ray3.14 | Reasoning model, native 16-bit HDR, fast/cheap drafts | Partial | up to ~10s · 1080p | low | Luma API (lumaai), aggregators |
| Kling 3.0 | Physics, multi-shot consistency, up to 4K@60fps | Yes | 3–15s (up to ~2 min) | ~$0.09–0.14 | Aggregators (Replicate, fal) |
| Seedance 2.5 | Cost leader, strong image-to-video | Yes | short · up to 1080p | low | Aggregators (fal, Replicate) |
| Pika | Stylised, effects, social | Partial | short | low | Aggregators |
| OpenAI Sora 2 ⚠️ | Coherent physics, prompt fidelity | Yes | 4/8/12s (Pro 10/15/25s) | ~$0.10 base · $0.30–0.70 Pro | API discontinuing 2026-09-24 — do not build on it |
⚠️ Sora is exiting. OpenAI notified developers on 2026-03-24 that the Videos API and the
sora-2/sora-2-promodels (and their dated snapshots) are deprecated and shut down on 2026-09-24; the consumer app was already discontinued 2026-04-26 and OpenAI has announced no successor. Do not start new work on Sora — reach for Veo, Kling, or an aggregator instead.
Selection heuristic: prototype on a cheap model through an aggregator (one SDK, swap the model id string), then promote the specific clip to Veo 3.1 only when it ships to a client. Reach for first-party Gemini/Vertex when you need Google's enterprise SLA, data-residency, or 4K.
Official Documentation
| Source | URL | What it covers |
|---|---|---|
| fal queue API | https://fal.ai/docs/model-apis/model-endpoints/queue | fal.queue.submit/status/result, subscribe, webhookUrl |
| fal Veo 3.1 endpoint | https://fal.ai/models/fal-ai/veo3.1/api | Input schema, duration/resolution/aspect enums (the old fal-ai/veo3 endpoint is deprecated) |
| Replicate Veo 3.1 | https://replicate.com/google/veo-3.1 | Model id, image-to-video, durations (Fast = google/veo-3.1-fast) |
| Replicate JS client | https://github.com/replicate/replicate-javascript | predictions.create + webhook, predictions.get, run |
| Gemini Veo docs | https://ai.google.dev/gemini-api/docs/veo | @google/genai generateVideos, operation polling, download |
| Luma API | https://docs.lumalabs.ai/docs/api | Create generation → id → poll status (Ray3 / Ray3.14; Luma Agents API for Ray3.2) |
| Runway Dev API | https://docs.dev.runwayml.com/ | Unified API (Gen-4.5, Gen-4 Turbo, Aleph 2.0, Act-Two); changelog for deprecations |
The async lifecycle: submit → poll / webhook → download
Every provider is a variation on this. Submit returns an id; the clip is not in that response. You then wait by polling or by receiving a webhook, then download.
sequenceDiagram
participant App as Your app
participant V as Video provider
participant CB as Your webhook (optional)
participant R2 as R2 / Supabase
App->>V: POST prompt (+ image) · submit job
V-->>App: 202 · { job_id, status: "queued" }
alt Polling
loop every few seconds
App->>V: GET job_id
V-->>App: status: queued / processing / succeeded
end
else Webhook
V->>CB: POST job_id · status: succeeded · video url
CB->>CB: verify · dedupe on job_id
end
App->>V: GET signed clip url
App->>R2: re-host MP4 (provider url expires)
Two render paths, one rule: the clip URL the provider hands back is temporary (Veo deletes after ~2 days; aggregator URLs are signed and expire). Download and re-host to your own bucket immediately — never store the provider URL as your permanent asset link.
Provider selection flow
flowchart TD
A["Need a video"] --> B{"Have a start image?"}
B -->|"yes"| C["Image-to-video"]
B -->|"no"| D["Text-to-video"]
C --> E{"Budget?"}
D --> E
E -->|"draft / social · cheap"| F["Kling / Luma / Veo-Fast via aggregator"]
E -->|"client deliverable · quality"| G["Veo 3.1 Standard / Runway Gen-4.5"]
F --> H{"Scale / many jobs?"}
G --> H
H -->|"yes"| I["Use webhookUrl · no polling"]
H -->|"no · one-off"| J["subscribe / poll inline"]
fal.ai — the cleanest aggregator (submit + webhook OR subscribe)
@fal-ai/client exposes the queue directly. For production, submit with a webhookUrl so you never block. For a quick script, fal.subscribe hides the polling.
// lib/video/fal.ts
import { fal } from "@fal-ai/client";
fal.config({ credentials: process.env.FAL_KEY! }); // server-side only
/** Production path: submit and let fal call your webhook when done. */
export async function submitVeoJob(prompt: string): Promise<string> {
// fal-ai/veo3.1 is the current endpoint — the old fal-ai/veo3 is deprecated.
const { request_id } = await fal.queue.submit("fal-ai/veo3.1", {
input: {
prompt,
aspect_ratio: "9:16", // vertical for social / WhatsApp status
duration: "8s", // string enum: "4s" | "6s" | "8s"
resolution: "720p", // draft res; "1080p" / "4k" require duration "8s"
generate_audio: true, // default true — audio is native to Veo 3.1
},
webhookUrl: "https://app.codeamanilabs.org/api/webhooks/fal",
});
return request_id; // store this — it's your idempotency key on the callback
}
/** Read the finished clip once the webhook says it's ready. */
export async function fetchVeoResult(requestId: string): Promise<string> {
const result = await fal.queue.result("fal-ai/veo3.1", { requestId });
return result.data.video.url; // re-host this immediately (it expires)
}
For a one-off (no webhook infrastructure), subscribe polls for you:
const result = await fal.subscribe("fal-ai/veo3.1", {
input: { prompt, duration: "8s", resolution: "720p" },
logs: true,
});
// result.data.video.url
The webhook receiver is exactly the pattern in webhooks/CLAUDE_CODE_INTEGRATION.md: verify, dedupe on request_id, ACK fast, then download + re-host out of band.
Replicate — one client, hundreds of models
replicate.predictions.create returns a prediction id and supports webhook + webhook_events_filter. Polling is predictions.get(id) with statuses starting → processing → succeeded | failed.
// lib/video/replicate.ts
import Replicate from "replicate";
const replicate = new Replicate({ auth: process.env.REPLICATE_API_TOKEN! });
/** Submit a Veo 3.1 image-to-video render with a webhook callback. */
export async function submitReplicateVideo(prompt: string, imageUrl?: string) {
const prediction = await replicate.predictions.create({
model: "google/veo-3.1",
input: {
prompt,
image: imageUrl, // omit for text-to-video
duration: 8, // 4 | 6 | 8 seconds
resolution: "1080p", // 720p | 1080p @ 24fps
aspect_ratio: "9:16",
},
webhook: "https://app.codeamanilabs.org/api/webhooks/replicate",
webhook_events_filter: ["completed"], // fire once, when terminal
});
return prediction.id;
}
/** Polling fallback when you have no webhook endpoint. */
export async function pollReplicate(id: string): Promise<string> {
let prediction = await replicate.predictions.get(id);
while (prediction.status !== "succeeded" && prediction.status !== "failed") {
await new Promise((r) => setTimeout(r, 3000));
prediction = await replicate.predictions.get(id);
}
if (prediction.status === "failed") throw new Error(String(prediction.error));
return prediction.output as unknown as string; // signed MP4 url → re-host
}
replicate.run(model, { input }) is the blocking convenience form — fine for scripts, avoid in request handlers because a render can outlast a serverless function's timeout.
Google Veo via the Gemini API (@google/genai) — first-party
The first-party path uses a long-running operation: generateVideos returns an operation you poll with getVideosOperation until operation.done, then download via ai.files.download.
// lib/video/veo.ts
import { GoogleGenAI } from "@google/genai";
const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY! });
export async function generateVeoClip(prompt: string): Promise<string> {
let operation = await ai.models.generateVideos({
// veo-3.1-generate-preview (standard) · veo-3.1-fast-generate-preview (drafts)
// · veo-3.1-lite-generate-preview (cheapest, no 4K / no reference images)
model: "veo-3.1-generate-preview",
prompt,
});
// Poll the long-running operation (no inline webhook on this SDK path).
while (!operation.done) {
await new Promise((r) => setTimeout(r, 10_000));
operation = await ai.operations.getVideosOperation({ operation });
}
const video = operation.response.generatedVideos[0].video;
await ai.files.download({ file: video, downloadPath: "output.mp4" });
return "output.mp4"; // upload to R2/Supabase — Gemini deletes after ~2 days
}
Veo constraints: duration is 4 / 6 / 8s; 8s is required for 1080p/4K or when supplying reference images; output includes natively generated audio. For background jobs at scale on Google, prefer Vertex AI (it exposes proper async operations + Cloud Storage output) over the inline polling loop.
Cost, limits & content moderation
- Cost discipline. Price is per output-second and renders are not free to retry. Draft at 720p on a cheap model, deliver at 1080p on Veo 3.1 / Runway. Cap
duration(a 4s clip is half the cost of 8s). Show users an estimated cost before they hit generate. - Duration/resolution caps are enums, not free numbers. Most models only accept fixed steps (4/6/8s; 720p/1080p; 16:9 or 9:16). Validate the user's choice against the enum at the boundary — an invalid value is a wasted round-trip.
- Prompt structure matters. Treat the prompt as a shot list: subject + action + setting + camera move + lighting + style. e.g. "Medium shot, a Nairobi street vendor arranging mangoes at dawn, slow dolly-in, warm golden light, cinematic." Use
negative_promptto suppress artefacts; supply a start image for brand/character consistency (image-to-video). - Content moderation & safety. Every provider runs input + output safety filters (Veo's
safety_tolerance, fal'ssafety_tolerance, Runway's moderation). Real-person likeness, public figures, and certain content are blocked — a job can come backfailed/moderated, so handle that branch and don't bill the user for a rejected render. Note that Veo (and most models) stamp output with an invisible SynthID watermark, so AI-generated provenance travels with the clip. - Idempotency. Store the
job_idthe instantsubmitreturns. On webhook delivery, dedupe on it (at-least-once delivery — same rules aswebhooks/).
codeAmani notes
-
New row in the AI routing policy. Video joins Claude (reasoning) / OpenAI (structured) / HuggingFace (open) as a distinct modality:
Use case Provider Client-deliverable / hero video, lip-sync, 1080p Veo 3.1 (Gemini API or Vertex) Draft / social / WhatsApp-status clips, cost-first Kling / Luma / Veo-Fast via aggregator Swap-models-fast prototyping Replicate or fal.ai (one SDK, change the id string) -
East African marketing content. AI video is a force-multiplier for SME social marketing — product reels, M-Pesa promo clips, WhatsApp-status ads. Default to 9:16 vertical (the WhatsApp/TikTok/Reels format dominant on Android here) and keep clips short (4–6s) to stay cheap and to load on 2G/3G.
-
Swahili prompts. The strong models accept Swahili/Sheng prompts directly ("Muuzaji wa maembe sokoni Nairobi, mwanga wa asubuhi, cinematic"). For best fidelity, write the visual description in English but keep any on-screen text / dialogue in Swahili — and verify rendered captions, since non-Latin/loanword spelling can drift.
-
Cost discipline on a budget. Never render in a request handler — always submit + webhook, store the
job_id, and re-host the finished MP4 to R2 (cheap egress) rather than serving the provider's expiring URL. Draft cheap, deliver premium, cap duration, and surface the per-clip cost estimate to the user before generating. -
Secrets.
FAL_KEY,REPLICATE_API_TOKEN,GEMINI_API_KEYare server-side only —.env.locallocally, Vercel env vars in prod, neverNEXT_PUBLIC_*. Verify the provider's webhook signature before trusting the callback (seewebhooks/). -
Route layout.
app/api/webhooks/<provider>/route.tsfor fal/Replicate render callbacks; render-submit logic lives inlib/video/<provider>.ts.
Official docs: