← Back to dashboard

AI Video Generation Integration Guide

What is AI video generation?

The real model

An async job pipeline: submit → poll/webhook → download → re-host. Priced per second.

Call models first-party (Veo via @google/genai) or through aggregators (Replicate google/veo-3.1, fal.ai fal-ai/veo3.1) that put dozens behind one SDK — swap the model id string to change models. The render is long-running, so always submit + webhook in production (the same receiver pattern the repo documents for webhooks); never block a serverless handler. Cost is per output-second, so draft cheap at 720p and only promote to Veo 3.1 / Runway Gen-4.5 for delivery. For codeAmani this is a new modality in the AI routing policy: vertical 9:16 social clips for the East African market, Swahili prompts, cost capped by short durations.

Six things to know about generative video

It's an async render, not a request/response — and the clip URL is temporary.

Text
 █████╗ ██╗    ██╗   ██╗██╗██████╗ ███████╗ ██████╗
██╔══██╗██║    ██║   ██║██║██╔══██╗██╔════╝██╔═══██╗
███████║██║    ██║   ██║██║██║  ██║█████╗  ██║   ██║
██╔══██║██║    ╚██╗ ██╔╝██║██║  ██║██╔══╝  ██║   ██║
██║  ██║██║     ╚████╔╝ ██║██████╔╝███████╗╚██████╔╝
╚═╝  ╚═╝╚═╝      ╚═══╝  ╚═╝╚═════╝ ╚══════╝ ╚═════╝

AI Video Generation Integration Guide

Focus: generative video is an async job, not a request/response. You submit a prompt, get a job id, then either poll the status or let a webhook call you back, and finally download the rendered clip. It is the exact reverse-API flow the repo already documents for webhooks — applied to a render that takes 30 seconds to several minutes.

Overview

Text-to-video and image-to-video models turn a prompt (and optionally a starting image) into a short MP4 with — increasingly — native audio. Unlike an LLM call that streams tokens back in seconds, a video render is long-running and expensive: a single 8-second 1080p clip can take a minute or two of GPU time and cost anywhere from a few cents to several dollars. Because of that, no serious provider makes you hold an HTTP connection open for the render. They all converged on the same shape:

  1. Submit the prompt → you immediately get a job / prediction / operation id.
  2. Wait for completion via one of two mechanisms:
    • Poll — GET the job id every few seconds until status is terminal.
    • Webhook — register a callback URL; the provider POSTs you when the clip is ready (this is the cheaper, more scalable path, and it is the same receiver pattern documented in webhooks/).
  3. Download the rendered clip from the signed URL in the result (the file usually expires — re-host it to your own R2/Supabase bucket).

This guide treats AI video as a new modality in the codeAmani AI routing policy — alongside Claude (reasoning), OpenAI (structured), and HuggingFace (open models) — and shows how to call it in production through both first-party APIs (Google Veo via the Gemini API) and aggregators (Replicate, fal.ai) that put dozens of models behind one async interface.

Models & providers

Prices are per second of output and move fast — always re-check the model page before committing a budget. As of the review date (2026-08-23):

ModelBest atNative audioDuration / res~Cost/secHow to call
Google Veo 3.1Top-tier realism, lip-sync, 1080p/4KYes (always on)4/6/8s · 720p/1080p/4K~$0.40 (Standard); ~$0.15 FastGemini API (@google/genai), Vertex AI, Replicate (google/veo-3.1), fal (fal-ai/veo3.1)
Runway Gen-4.5Cinematic control, native 4K, prompt adherenceNo (Aleph 2.0 edits w/ audio)short clips · up to 4K~$0.20Runway Dev API (@runwayml/sdk), aggregators
Luma Ray3 / Ray3.14Reasoning model, native 16-bit HDR, fast/cheap draftsPartialup to ~10s · 1080plowLuma API (lumaai), aggregators
Kling 3.0Physics, multi-shot consistency, up to 4K@60fpsYes3–15s (up to ~2 min)~$0.09–0.14Aggregators (Replicate, fal)
Seedance 2.5Cost leader, strong image-to-videoYesshort · up to 1080plowAggregators (fal, Replicate)
PikaStylised, effects, socialPartialshortlowAggregators
OpenAI Sora 2 ⚠️Coherent physics, prompt fidelityYes4/8/12s (Pro 10/15/25s)~$0.10 base · $0.30–0.70 ProAPI discontinuing 2026-09-24 — do not build on it

⚠️ Sora is exiting. OpenAI notified developers on 2026-03-24 that the Videos API and the sora-2 / sora-2-pro models (and their dated snapshots) are deprecated and shut down on 2026-09-24; the consumer app was already discontinued 2026-04-26 and OpenAI has announced no successor. Do not start new work on Sora — reach for Veo, Kling, or an aggregator instead.

Selection heuristic: prototype on a cheap model through an aggregator (one SDK, swap the model id string), then promote the specific clip to Veo 3.1 only when it ships to a client. Reach for first-party Gemini/Vertex when you need Google's enterprise SLA, data-residency, or 4K.

Official Documentation

SourceURLWhat it covers
fal queue APIhttps://fal.ai/docs/model-apis/model-endpoints/queuefal.queue.submit/status/result, subscribe, webhookUrl
fal Veo 3.1 endpointhttps://fal.ai/models/fal-ai/veo3.1/apiInput schema, duration/resolution/aspect enums (the old fal-ai/veo3 endpoint is deprecated)
Replicate Veo 3.1https://replicate.com/google/veo-3.1Model id, image-to-video, durations (Fast = google/veo-3.1-fast)
Replicate JS clienthttps://github.com/replicate/replicate-javascriptpredictions.create + webhook, predictions.get, run
Gemini Veo docshttps://ai.google.dev/gemini-api/docs/veo@google/genai generateVideos, operation polling, download
Luma APIhttps://docs.lumalabs.ai/docs/apiCreate generation → id → poll status (Ray3 / Ray3.14; Luma Agents API for Ray3.2)
Runway Dev APIhttps://docs.dev.runwayml.com/Unified API (Gen-4.5, Gen-4 Turbo, Aleph 2.0, Act-Two); changelog for deprecations

The async lifecycle: submit → poll / webhook → download

Every provider is a variation on this. Submit returns an id; the clip is not in that response. You then wait by polling or by receiving a webhook, then download.

Two render paths, one rule: the clip URL the provider hands back is temporary (Veo deletes after ~2 days; aggregator URLs are signed and expire). Download and re-host to your own bucket immediately — never store the provider URL as your permanent asset link.

Provider selection flow


fal.ai — the cleanest aggregator (submit + webhook OR subscribe)

@fal-ai/client exposes the queue directly. For production, submit with a webhookUrl so you never block. For a quick script, fal.subscribe hides the polling.

TypeScript
// lib/video/fal.ts
import { fal } from "@fal-ai/client";

fal.config({ credentials: process.env.FAL_KEY! }); // server-side only

/** Production path: submit and let fal call your webhook when done. */
export async function submitVeoJob(prompt: string): Promise<string> {
  // fal-ai/veo3.1 is the current endpoint — the old fal-ai/veo3 is deprecated.
  const { request_id } = await fal.queue.submit("fal-ai/veo3.1", {
    input: {
      prompt,
      aspect_ratio: "9:16", // vertical for social / WhatsApp status
      duration: "8s",       // string enum: "4s" | "6s" | "8s"
      resolution: "720p",   // draft res; "1080p" / "4k" require duration "8s"
      generate_audio: true, // default true — audio is native to Veo 3.1
    },
    webhookUrl: "https://app.codeamanilabs.org/api/webhooks/fal",
  });
  return request_id; // store this — it's your idempotency key on the callback
}

/** Read the finished clip once the webhook says it's ready. */
export async function fetchVeoResult(requestId: string): Promise<string> {
  const result = await fal.queue.result("fal-ai/veo3.1", { requestId });
  return result.data.video.url; // re-host this immediately (it expires)
}

For a one-off (no webhook infrastructure), subscribe polls for you:

TypeScript
const result = await fal.subscribe("fal-ai/veo3.1", {
  input: { prompt, duration: "8s", resolution: "720p" },
  logs: true,
});
// result.data.video.url

The webhook receiver is exactly the pattern in webhooks/CLAUDE_CODE_INTEGRATION.md: verify, dedupe on request_id, ACK fast, then download + re-host out of band.


Replicate — one client, hundreds of models

replicate.predictions.create returns a prediction id and supports webhook + webhook_events_filter. Polling is predictions.get(id) with statuses starting → processing → succeeded | failed.

TypeScript
// lib/video/replicate.ts
import Replicate from "replicate";

const replicate = new Replicate({ auth: process.env.REPLICATE_API_TOKEN! });

/** Submit a Veo 3.1 image-to-video render with a webhook callback. */
export async function submitReplicateVideo(prompt: string, imageUrl?: string) {
  const prediction = await replicate.predictions.create({
    model: "google/veo-3.1",
    input: {
      prompt,
      image: imageUrl,        // omit for text-to-video
      duration: 8,            // 4 | 6 | 8 seconds
      resolution: "1080p",    // 720p | 1080p @ 24fps
      aspect_ratio: "9:16",
    },
    webhook: "https://app.codeamanilabs.org/api/webhooks/replicate",
    webhook_events_filter: ["completed"], // fire once, when terminal
  });
  return prediction.id;
}

/** Polling fallback when you have no webhook endpoint. */
export async function pollReplicate(id: string): Promise<string> {
  let prediction = await replicate.predictions.get(id);
  while (prediction.status !== "succeeded" && prediction.status !== "failed") {
    await new Promise((r) => setTimeout(r, 3000));
    prediction = await replicate.predictions.get(id);
  }
  if (prediction.status === "failed") throw new Error(String(prediction.error));
  return prediction.output as unknown as string; // signed MP4 url → re-host
}

replicate.run(model, { input }) is the blocking convenience form — fine for scripts, avoid in request handlers because a render can outlast a serverless function's timeout.


Google Veo via the Gemini API (@google/genai) — first-party

The first-party path uses a long-running operation: generateVideos returns an operation you poll with getVideosOperation until operation.done, then download via ai.files.download.

TypeScript
// lib/video/veo.ts
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY! });

export async function generateVeoClip(prompt: string): Promise<string> {
  let operation = await ai.models.generateVideos({
    // veo-3.1-generate-preview (standard) · veo-3.1-fast-generate-preview (drafts)
    // · veo-3.1-lite-generate-preview (cheapest, no 4K / no reference images)
    model: "veo-3.1-generate-preview",
    prompt,
  });

  // Poll the long-running operation (no inline webhook on this SDK path).
  while (!operation.done) {
    await new Promise((r) => setTimeout(r, 10_000));
    operation = await ai.operations.getVideosOperation({ operation });
  }

  const video = operation.response.generatedVideos[0].video;
  await ai.files.download({ file: video, downloadPath: "output.mp4" });
  return "output.mp4"; // upload to R2/Supabase — Gemini deletes after ~2 days
}

Veo constraints: duration is 4 / 6 / 8s; 8s is required for 1080p/4K or when supplying reference images; output includes natively generated audio. For background jobs at scale on Google, prefer Vertex AI (it exposes proper async operations + Cloud Storage output) over the inline polling loop.


Cost, limits & content moderation

  • Cost discipline. Price is per output-second and renders are not free to retry. Draft at 720p on a cheap model, deliver at 1080p on Veo 3.1 / Runway. Cap duration (a 4s clip is half the cost of 8s). Show users an estimated cost before they hit generate.
  • Duration/resolution caps are enums, not free numbers. Most models only accept fixed steps (4/6/8s; 720p/1080p; 16:9 or 9:16). Validate the user's choice against the enum at the boundary — an invalid value is a wasted round-trip.
  • Prompt structure matters. Treat the prompt as a shot list: subject + action + setting + camera move + lighting + style. e.g. "Medium shot, a Nairobi street vendor arranging mangoes at dawn, slow dolly-in, warm golden light, cinematic." Use negative_prompt to suppress artefacts; supply a start image for brand/character consistency (image-to-video).
  • Content moderation & safety. Every provider runs input + output safety filters (Veo's safety_tolerance, fal's safety_tolerance, Runway's moderation). Real-person likeness, public figures, and certain content are blocked — a job can come back failed/moderated, so handle that branch and don't bill the user for a rejected render. Note that Veo (and most models) stamp output with an invisible SynthID watermark, so AI-generated provenance travels with the clip.
  • Idempotency. Store the job_id the instant submit returns. On webhook delivery, dedupe on it (at-least-once delivery — same rules as webhooks/).

codeAmani notes

  • New row in the AI routing policy. Video joins Claude (reasoning) / OpenAI (structured) / HuggingFace (open) as a distinct modality:

    Use caseProvider
    Client-deliverable / hero video, lip-sync, 1080pVeo 3.1 (Gemini API or Vertex)
    Draft / social / WhatsApp-status clips, cost-firstKling / Luma / Veo-Fast via aggregator
    Swap-models-fast prototypingReplicate or fal.ai (one SDK, change the id string)
  • East African marketing content. AI video is a force-multiplier for SME social marketing — product reels, M-Pesa promo clips, WhatsApp-status ads. Default to 9:16 vertical (the WhatsApp/TikTok/Reels format dominant on Android here) and keep clips short (4–6s) to stay cheap and to load on 2G/3G.

  • Swahili prompts. The strong models accept Swahili/Sheng prompts directly ("Muuzaji wa maembe sokoni Nairobi, mwanga wa asubuhi, cinematic"). For best fidelity, write the visual description in English but keep any on-screen text / dialogue in Swahili — and verify rendered captions, since non-Latin/loanword spelling can drift.

  • Cost discipline on a budget. Never render in a request handler — always submit + webhook, store the job_id, and re-host the finished MP4 to R2 (cheap egress) rather than serving the provider's expiring URL. Draft cheap, deliver premium, cap duration, and surface the per-clip cost estimate to the user before generating.

  • Secrets. FAL_KEY, REPLICATE_API_TOKEN, GEMINI_API_KEY are server-side only — .env.local locally, Vercel env vars in prod, never NEXT_PUBLIC_*. Verify the provider's webhook signature before trusting the callback (see webhooks/).

  • Route layout. app/api/webhooks/<provider>/route.ts for fal/Replicate render callbacks; render-submit logic lives in lib/video/<provider>.ts.