AI Video Generation Integration Guide

Technology: ai-video · Category: ai · Last reviewed: 2026-08-23

Source: https://tech-stack.codeamanilabs.org/guide/ai-video

Insight:

Generative video is an async job, not a request/response: you submit a prompt, get a job id, then poll or receive a webhook before downloading a clip that costs cents-to-dollars per output-second and whose URL expires within days. The market moved fast through 2026 — Veo 3.1 is the quality leader, aggregators (fal fal-ai/veo3.1, Replicate google/veo-3.1) let you swap models with one string, and OpenAI is exiting: the Sora 2 API shuts down 2026-09-24 with no announced successor. For codeAmani it is a distinct AI-routing modality — default to short 9:16 vertical clips for East African social / WhatsApp marketing, draft cheap on a fast tier, and re-host every render to R2 before the provider URL expires.

 █████╗ ██╗    ██╗   ██╗██╗██████╗ ███████╗ ██████╗
██╔══██╗██║    ██║   ██║██║██╔══██╗██╔════╝██╔═══██╗
███████║██║    ██║   ██║██║██║  ██║█████╗  ██║   ██║
██╔══██║██║    ╚██╗ ██╔╝██║██║  ██║██╔══╝  ██║   ██║
██║  ██║██║     ╚████╔╝ ██║██████╔╝███████╗╚██████╔╝
╚═╝  ╚═╝╚═╝      ╚═══╝  ╚═╝╚═════╝ ╚══════╝ ╚═════╝

AI Video Generation Integration Guide

Focus: generative video is an async job, not a request/response. You submit a prompt, get a job id, then either poll the status or let a webhook call you back, and finally download the rendered clip. It is the exact reverse-API flow the repo already documents for webhooks — applied to a render that takes 30 seconds to several minutes.

Overview

Text-to-video and image-to-video models turn a prompt (and optionally a starting image) into a short MP4 with — increasingly — native audio. Unlike an LLM call that streams tokens back in seconds, a video render is long-running and expensive: a single 8-second 1080p clip can take a minute or two of GPU time and cost anywhere from a few cents to several dollars. Because of that, no serious provider makes you hold an HTTP connection open for the render. They all converged on the same shape:

  1. Submit the prompt → you immediately get a job / prediction / operation id.
  2. Wait for completion via one of two mechanisms:
    • Poll — GET the job id every few seconds until status is terminal.
    • Webhook — register a callback URL; the provider POSTs you when the clip is ready (this is the cheaper, more scalable path, and it is the same receiver pattern documented in webhooks/).
  3. Download the rendered clip from the signed URL in the result (the file usually expires — re-host it to your own R2/Supabase bucket).

This guide treats AI video as a new modality in the codeAmani AI routing policy — alongside Claude (reasoning), OpenAI (structured), and HuggingFace (open models) — and shows how to call it in production through both first-party APIs (Google Veo via the Gemini API) and aggregators (Replicate, fal.ai) that put dozens of models behind one async interface.

Models & providers

Prices are per second of output and move fast — always re-check the model page before committing a budget. As of the review date (2026-08-23):

Model Best at Native audio Duration / res ~Cost/sec How to call
Google Veo 3.1 Top-tier realism, lip-sync, 1080p/4K Yes (always on) 4/6/8s · 720p/1080p/4K ~$0.40 (Standard); ~$0.15 Fast Gemini API (@google/genai), Vertex AI, Replicate (google/veo-3.1), fal (fal-ai/veo3.1)
Runway Gen-4.5 Cinematic control, native 4K, prompt adherence No (Aleph 2.0 edits w/ audio) short clips · up to 4K ~$0.20 Runway Dev API (@runwayml/sdk), aggregators
Luma Ray3 / Ray3.14 Reasoning model, native 16-bit HDR, fast/cheap drafts Partial up to ~10s · 1080p low Luma API (lumaai), aggregators
Kling 3.0 Physics, multi-shot consistency, up to 4K@60fps Yes 3–15s (up to ~2 min) ~$0.09–0.14 Aggregators (Replicate, fal)
Seedance 2.5 Cost leader, strong image-to-video Yes short · up to 1080p low Aggregators (fal, Replicate)
Pika Stylised, effects, social Partial short low Aggregators
OpenAI Sora 2 ⚠️ Coherent physics, prompt fidelity Yes 4/8/12s (Pro 10/15/25s) ~$0.10 base · $0.30–0.70 Pro API discontinuing 2026-09-24 — do not build on it

⚠️ Sora is exiting. OpenAI notified developers on 2026-03-24 that the Videos API and the sora-2 / sora-2-pro models (and their dated snapshots) are deprecated and shut down on 2026-09-24; the consumer app was already discontinued 2026-04-26 and OpenAI has announced no successor. Do not start new work on Sora — reach for Veo, Kling, or an aggregator instead.

Selection heuristic: prototype on a cheap model through an aggregator (one SDK, swap the model id string), then promote the specific clip to Veo 3.1 only when it ships to a client. Reach for first-party Gemini/Vertex when you need Google's enterprise SLA, data-residency, or 4K.

Official Documentation

Source URL What it covers
fal queue API https://fal.ai/docs/model-apis/model-endpoints/queue fal.queue.submit/status/result, subscribe, webhookUrl
fal Veo 3.1 endpoint https://fal.ai/models/fal-ai/veo3.1/api Input schema, duration/resolution/aspect enums (the old fal-ai/veo3 endpoint is deprecated)
Replicate Veo 3.1 https://replicate.com/google/veo-3.1 Model id, image-to-video, durations (Fast = google/veo-3.1-fast)
Replicate JS client https://github.com/replicate/replicate-javascript predictions.create + webhook, predictions.get, run
Gemini Veo docs https://ai.google.dev/gemini-api/docs/veo @google/genai generateVideos, operation polling, download
Luma API https://docs.lumalabs.ai/docs/api Create generation → id → poll status (Ray3 / Ray3.14; Luma Agents API for Ray3.2)
Runway Dev API https://docs.dev.runwayml.com/ Unified API (Gen-4.5, Gen-4 Turbo, Aleph 2.0, Act-Two); changelog for deprecations

The async lifecycle: submit → poll / webhook → download

Every provider is a variation on this. Submit returns an id; the clip is not in that response. You then wait by polling or by receiving a webhook, then download.

sequenceDiagram
    participant App as Your app
    participant V as Video provider
    participant CB as Your webhook (optional)
    participant R2 as R2 / Supabase
    App->>V: POST prompt (+ image) · submit job
    V-->>App: 202 · { job_id, status: "queued" }
    alt Polling
        loop every few seconds
            App->>V: GET job_id
            V-->>App: status: queued / processing / succeeded
        end
    else Webhook
        V->>CB: POST job_id · status: succeeded · video url
        CB->>CB: verify · dedupe on job_id
    end
    App->>V: GET signed clip url
    App->>R2: re-host MP4 (provider url expires)

Two render paths, one rule: the clip URL the provider hands back is temporary (Veo deletes after ~2 days; aggregator URLs are signed and expire). Download and re-host to your own bucket immediately — never store the provider URL as your permanent asset link.

Provider selection flow

flowchart TD
    A["Need a video"] --> B{"Have a start image?"}
    B -->|"yes"| C["Image-to-video"]
    B -->|"no"| D["Text-to-video"]
    C --> E{"Budget?"}
    D --> E
    E -->|"draft / social · cheap"| F["Kling / Luma / Veo-Fast via aggregator"]
    E -->|"client deliverable · quality"| G["Veo 3.1 Standard / Runway Gen-4.5"]
    F --> H{"Scale / many jobs?"}
    G --> H
    H -->|"yes"| I["Use webhookUrl · no polling"]
    H -->|"no · one-off"| J["subscribe / poll inline"]

fal.ai — the cleanest aggregator (submit + webhook OR subscribe)

@fal-ai/client exposes the queue directly. For production, submit with a webhookUrl so you never block. For a quick script, fal.subscribe hides the polling.

// lib/video/fal.ts
import { fal } from "@fal-ai/client";

fal.config({ credentials: process.env.FAL_KEY! }); // server-side only

/** Production path: submit and let fal call your webhook when done. */
export async function submitVeoJob(prompt: string): Promise<string> {
  // fal-ai/veo3.1 is the current endpoint — the old fal-ai/veo3 is deprecated.
  const { request_id } = await fal.queue.submit("fal-ai/veo3.1", {
    input: {
      prompt,
      aspect_ratio: "9:16", // vertical for social / WhatsApp status
      duration: "8s",       // string enum: "4s" | "6s" | "8s"
      resolution: "720p",   // draft res; "1080p" / "4k" require duration "8s"
      generate_audio: true, // default true — audio is native to Veo 3.1
    },
    webhookUrl: "https://app.codeamanilabs.org/api/webhooks/fal",
  });
  return request_id; // store this — it's your idempotency key on the callback
}

/** Read the finished clip once the webhook says it's ready. */
export async function fetchVeoResult(requestId: string): Promise<string> {
  const result = await fal.queue.result("fal-ai/veo3.1", { requestId });
  return result.data.video.url; // re-host this immediately (it expires)
}

For a one-off (no webhook infrastructure), subscribe polls for you:

const result = await fal.subscribe("fal-ai/veo3.1", {
  input: { prompt, duration: "8s", resolution: "720p" },
  logs: true,
});
// result.data.video.url

The webhook receiver is exactly the pattern in webhooks/CLAUDE_CODE_INTEGRATION.md: verify, dedupe on request_id, ACK fast, then download + re-host out of band.


Replicate — one client, hundreds of models

replicate.predictions.create returns a prediction id and supports webhook + webhook_events_filter. Polling is predictions.get(id) with statuses starting → processing → succeeded | failed.

// lib/video/replicate.ts
import Replicate from "replicate";

const replicate = new Replicate({ auth: process.env.REPLICATE_API_TOKEN! });

/** Submit a Veo 3.1 image-to-video render with a webhook callback. */
export async function submitReplicateVideo(prompt: string, imageUrl?: string) {
  const prediction = await replicate.predictions.create({
    model: "google/veo-3.1",
    input: {
      prompt,
      image: imageUrl,        // omit for text-to-video
      duration: 8,            // 4 | 6 | 8 seconds
      resolution: "1080p",    // 720p | 1080p @ 24fps
      aspect_ratio: "9:16",
    },
    webhook: "https://app.codeamanilabs.org/api/webhooks/replicate",
    webhook_events_filter: ["completed"], // fire once, when terminal
  });
  return prediction.id;
}

/** Polling fallback when you have no webhook endpoint. */
export async function pollReplicate(id: string): Promise<string> {
  let prediction = await replicate.predictions.get(id);
  while (prediction.status !== "succeeded" && prediction.status !== "failed") {
    await new Promise((r) => setTimeout(r, 3000));
    prediction = await replicate.predictions.get(id);
  }
  if (prediction.status === "failed") throw new Error(String(prediction.error));
  return prediction.output as unknown as string; // signed MP4 url → re-host
}

replicate.run(model, { input }) is the blocking convenience form — fine for scripts, avoid in request handlers because a render can outlast a serverless function's timeout.


Google Veo via the Gemini API (@google/genai) — first-party

The first-party path uses a long-running operation: generateVideos returns an operation you poll with getVideosOperation until operation.done, then download via ai.files.download.

// lib/video/veo.ts
import { GoogleGenAI } from "@google/genai";

const ai = new GoogleGenAI({ apiKey: process.env.GEMINI_API_KEY! });

export async function generateVeoClip(prompt: string): Promise<string> {
  let operation = await ai.models.generateVideos({
    // veo-3.1-generate-preview (standard) · veo-3.1-fast-generate-preview (drafts)
    // · veo-3.1-lite-generate-preview (cheapest, no 4K / no reference images)
    model: "veo-3.1-generate-preview",
    prompt,
  });

  // Poll the long-running operation (no inline webhook on this SDK path).
  while (!operation.done) {
    await new Promise((r) => setTimeout(r, 10_000));
    operation = await ai.operations.getVideosOperation({ operation });
  }

  const video = operation.response.generatedVideos[0].video;
  await ai.files.download({ file: video, downloadPath: "output.mp4" });
  return "output.mp4"; // upload to R2/Supabase — Gemini deletes after ~2 days
}

Veo constraints: duration is 4 / 6 / 8s; 8s is required for 1080p/4K or when supplying reference images; output includes natively generated audio. For background jobs at scale on Google, prefer Vertex AI (it exposes proper async operations + Cloud Storage output) over the inline polling loop.


Cost, limits & content moderation


codeAmani notes

Official docs: