This directory covers six current video-generation API providers and two APIs with announced lifecycle limits, using public documentation and prices checked July 30, 2026.
1. Google Vertex AI: Veo 3.1, Fast, and Lite
- Company and provider: Google provides Veo through Vertex AI in Google Cloud.
- Current models: The stable model IDs are
veo-3.1-generate-001andveo-3.1-fast-generate-001. Google listsveo-3.1-lite-generate-001as a preview model. The standard and Fast model pages list a November 17, 2025 release date and a retirement date of November 17, 2026 or later. The Lite page lists an April 2, 2026 preview release date. Earlierveo-3.1-generate-previewandveo-3.1-fast-generate-previewendpoints were removed April 2, 2026. - API endpoints, inputs, and outputs: Veo uses Vertex AI’s asynchronous video-generation API and returns a long-running operation record. The models accept text prompts and image inputs and return MP4 video. Standard, Fast, and Lite support text-to-video, image-to-video, first-and-last-frame generation, prompt rewriting, and video extension. Standard and Fast also list asset reference images. A request can ask for one to four outputs. Supported durations are 4, 6, or 8 seconds. Reference-image-to-video on the standard model uses an 8-second duration. The model pages list 16:9 and 9:16 aspect ratios, 24 fps output, and
video/mp4. Standard lists 720p, 1080p, and 4K output. Fast and Lite list 720p and 1080p in their technical specifications. The input image can be supplied from Cloud Storage or as supported request data, and completed output can be written to a Cloud Storage URI. - Documented controls: Request fields include prompt, source image or images for supported modes, first frame, last frame, output count, duration, aspect ratio, resolution, and prompt-rewriting behavior. Extension accepts a source video produced by Veo. Asset-reference mode supports reference images on the standard and Fast models but is not listed for Lite. Audio-capable price tiers generate video with synchronized speech, sound effects, or other audio. Sound generation is listed separately from the text and image input modalities on the model page. The current model pages list C2PA Content Credentials support. Vertex AI also exposes Cloud IAM, Cloud Storage destinations, customer-managed encryption keys, VPC Service Controls, and listed data-residency controls for supported configurations.
- Additional documented specifications: The standard model page lists text and image as input modalities and video as the output modality. It marks general video generation, extension, asset references, prompt rewriting, and C2PA as supported. Object insertion and object removal are listed as unsupported on that endpoint. Standard and Fast list general availability; Lite is preview. Each page publishes a model version table, release stage, release date, region, quota class, technical output table, and security-control table. Google publishes separate guides for image-to-video, first-and-last-frame generation, reference images, extension, prompt rewriting, and C2PA.
- Published price table: Vertex AI pricing bills successful output by generated second. Veo 3.1 Lite costs $0.03/sec for silent 720p, $0.05/sec for 720p with audio, $0.05/sec for silent 1080p, and $0.08/sec for 1080p with audio. Veo 3.1 Fast costs $0.08/sec for silent 720p, $0.10/sec for 720p with audio, $0.10/sec for silent 1080p, $0.12/sec for 1080p with audio, $0.25/sec for silent 4K, and $0.30/sec for 4K with audio. Standard Veo 3.1 costs $0.20/sec for silent 720p or 1080p, $0.40/sec for 720p or 1080p with audio, $0.40/sec for silent 4K, and $0.60/sec for 4K with audio. Google’s pricing page says requests that return a non-200 response are not charged for input or output.
- Documented job lifecycle and errors: A generation request returns a long-running operation. Vertex AI exposes operation retrieval until the job reaches a terminal result. Successful results contain output video references or write files to the configured Cloud Storage location. Request validation, quota exhaustion, content-policy filtering, and internal service failures are represented separately in API errors or operation results. Google’s general Vertex AI error documentation defines HTTP and RPC status handling, while the Veo guides define model-specific rejected parameter combinations.
- Limits: The image-to-video input limit is 20 MB. The listed prompt language is English. Input video is not listed as a general input modality; extension is a separate capability. Batch inference is unsupported. Standard and Fast list 50 regional online-prediction requests per base model per minute. The Lite page lists the same fixed-quota figure. The standard and Fast stable model pages list fixed quota and Provisioned Throughput rather than pay-as-you-go as consumption options. The pricing page includes a Fast 4K rate while the Fast technical-specification table lists 720p and 1080p output. The current pages list object insertion and removal as unsupported for the stable Veo 3.1 generation endpoints.
- Regions and availability: The current Veo 3.1 model pages list
us-central1. Standard and Fast are general availability. Lite is preview and is governed by Google Cloud’s pre-GA offering terms. - Published support and SLA terms: Veo supports fixed quota and Provisioned Throughput. Google’s Provisioned Throughput guide states that its enforcement window is separate from request latency. Provisioned Throughput uses purchased generation scale units and model-specific throughput calculations. Requests above purchased throughput can be handled according to the configured spillover behavior and available pay-as-you-go capacity. Google Cloud Customer Care offers Basic, Standard, Enhanced, and Premium plans. Standard and Enhanced plans publish response targets by case priority. Premium Support publishes a 15-minute initial response target for a P1 business-critical support case and includes named technical account management options. The public Veo pages do not publish a video-completion latency SLA.
- Data retention and training terms: Google’s service-specific terms state that Google will not use customer data to train or fine-tune generative AI models without prior permission or instruction. Vertex AI request and output storage depends on the customer’s configured Google Cloud resources and applicable data-processing terms.
- License and usage restrictions: Google Cloud’s service-specific terms contain generated-output and training-data indemnity provisions with eligibility conditions, exclusions, notice requirements, and remedies. Preview use is also subject to the pre-GA offering terms. Google’s acceptable-use and generative-AI prohibited-use policies apply to API requests and outputs.
- Provenance standards: The Veo 3.1 model pages list C2PA Content Credentials. Google also documents SynthID watermarking for generated media.
- Primary sources: Veo 3.1 model specifications, Vertex AI generative-model pricing, first-and-last-frame generation, Veo Provisioned Throughput, Google Cloud service-specific terms, and Google Cloud support.
2. Runway API: Gen-4.5, Gen-4 Turbo, and third-party models
- Company and provider: Runway AI, Inc. operates the Runway API and provides both Runway-developed models and selected third-party models through the service.
- Current models: The pricing catalog lists
gen4.5,gen4_turbo,aleph2,act_two, Seedance 2 variants, Veo variants, Gemini Omni Flash, animation, audio, and upscaling operations. Gen-4.5 supports text-to-video and image-to-video. Gen-4 Turbo is an image-to-video model. Aleph 2 performs video transformation. Act-Two transfers motion and performance from a driving video. The third-party entries retain their upstream model names. - API endpoints, inputs, and outputs: Runway uses asynchronous task endpoints and an explicit API-version request header. Version
2024-11-06is documented in the API-version section. The API reference includes text-to-video, image-to-video, video-to-video, character-performance, audio, and upscale operations. Inputs vary by model and can include text, an initial image, a source video, reference assets, or a driving-performance video. Video tasks return a task record followed by output URLs when processing completes. Task states include pending and running states,THROTTLED, success, and documented failure states. Runway also provides asset-upload and input-URL workflows, and the input guide specifies accepted media types and URL requirements. - Documented controls: Gen-4.5 accepts a seed and durations from 2 through 10 seconds. Its text-to-video output sizes include 1280×720 and 720×1280. The input-assets guide lists additional shapes for image-to-video and defines supported media inputs. Aleph 2 accepts a source video and transformation instructions. Act-Two accepts a character asset and driving performance. Available fields differ by endpoint and API version.
- Additional documented specifications: Runway publishes separate API sections for model routers, Characters, video recipes, image recipes, inputs, outputs, uploads, content moderation, usage tiers, attribution, organizations, and errors. Model Router response metadata identifies the model selected for the request and the credits charged. The public catalog includes generation, editing, performance transfer, upscale, audio, and real-time operations under one organization account. Runway’s version documentation identifies the request header and the API contract associated with version
2024-11-06. - Published price table: Runway pricing defines one credit as $0.01 and notes that sales tax can apply.
gen4.5costs 12 credits per output second.gen4_turboandact_twocost 5 credits per second.aleph2costs 28 credits per second with a 56-credit minimum.seedance2costs 36 credits/sec for 480p or 720p, 40 credits/sec for 1080p, and 150 credits/sec for 4K.seedance2_fastcosts 29 credits/sec for 480p or 720p.seedance2_minicosts 16 credits/sec for 480p or 720p and has a 64-credit minimum. Veo 3.1 costs 40 credits/sec with audio and 20 without audio. Veo 3.1 Fast costs 15 credits/sec with audio and 10 without audio.happyhorse_1_0costs 15 credits/sec at 720p and 30 at 1080p. Gemini Omni Flash costs 10 credits/sec for text-to-video and 10 credits/sec plus one credit for the initial image in image-to-video. Model Router generations are billed at the selected model’s standard rate, and response metadata reports the selected model and realized credit cost. - Documented job lifecycle and errors: Runway task creation returns an ID used to retrieve the task. Tasks beyond the active concurrency pool receive
THROTTLEDand later enter the execution queue. The task-failure guide separates invalid input or asset failures, safety or moderation outcomes, internal failures, and third-party-provider unavailability. Output URLs appear on successful task records. Runway’s HTTP error guide separately covers authentication, authorization, request format, version header, rate, and account errors. - Limits: Usage tiers define concurrency, daily-generation, and monthly-spend limits per organization. Tier 1 lists one concurrent video generation, 50 video generations per day, and $100 maximum monthly spend. Tiers 2 through 5 list video concurrency of 3, 5, 10, and 20. The same tiers list daily video-generation limits of 500, 1,000, 5,000, and 25,000 and monthly-spend limits of $500, $2,000, $20,000, and $100,000. Published upgrade criteria are $50 purchased for Tier 2, $100 for Tier 3, $1,000 for Tier 4, and $5,000 for Tier 5. All video-generation models share the organization’s video concurrency pool. Tasks submitted above the pool enter
THROTTLEDand remain stored until enqueued in approximate submission order. Runway states that there is no separate requests-per-minute maximum within the daily limit. - Regions and availability: The public API documentation does not publish a customer-selectable processing-region parameter for these endpoints. Model availability depends on the API catalog, account tier, enterprise agreement, and upstream-model availability.
- Published support and SLA terms: Higher limits, custom tiers, and guaranteed minimum concurrency are available through enterprise arrangements or a limits exception request. The usage-tier page states that actual concurrency can be below the listed maximum under load. The public self-serve documentation does not publish a universal task-completion SLA. The task-failure guide documents failures for invalid assets, internal processing, safety systems, and unavailable third-party providers.
- Data retention and training terms: Runway’s terms and privacy materials govern storage and processing for standard accounts. Runway’s enterprise third-party model FAQ states that enterprise customer data is not used for training under the described enterprise offering and describes contractual data-protection commitments for third-party models.
- License and usage restrictions: Runway’s terms state that Runway does not claim ownership of customer content. The standard API agreement includes “Powered by Runway” branding and end-user-terms requirements. Third-party models can carry upstream restrictions in addition to Runway’s terms. Enterprise terms and order forms can replace or supplement the standard provisions.
- Provenance standards: Provenance behavior is model-specific. Runway’s public API catalog does not state that every output from every model contains one common signed provenance format.
- Primary sources: Runway API reference, model pricing, usage tiers and queue behavior, input requirements, task failures, Runway terms, and enterprise third-party model FAQ.
Figure 1. Asynchronous video APIs expose queued and running states before a completed, failed, or moderated result. An unknown submission state is resolved by retrieving the existing job record.
3. Luma Agents API: Ray 3.2
- Company and provider: Luma AI, Inc. provides the Luma Agents API.
- Current models: The video model is
ray-3.2. The same API also listsuni-1anduni-1-maxfor image generation and image editing.ray-3.2supports thevideo,video_edit, andvideo_reframerequest types. - API endpoints, inputs, and outputs: The Agents API accepts text prompts, start frames, end frames, indexed keyframes, source video, prior Luma generation IDs, file URLs, uploaded files, and inline data where supported.
POST /v1/generationsis the shared generation endpoint described by the model and pricing guides.videocovers text-to-video and image-to-video.video_editaccepts a source identified by generation ID, URL, or data.video_reframeaccepts a source and target aspect ratio. Jobs are asynchronous and return a generation record followed by a presigned output URL. File inputs can be uploaded once through the Files API and referenced by file ID in later generation requests. - Documented controls: The Ray 3.2 model guide lists start and end frames, up to 64 keyframe anchors with explicit indexes, extension from prior generations, per-signal edit controls, reframing, loop mode, HDR output, and EXR export. Ordinary video generation accepts six aspect ratios. Video editing derives aspect ratio from the source and ignores an aspect-ratio request. Reframing requires a target ratio. HDR is available at 720p and 1080p. EXR export requires
hdr: true. HDR is unavailable for extension requests. - Additional documented specifications: The model matrix identifies
ray-3.2as the video model forvideo,video_edit, andvideo_reframe. Multi-keyframe requests pairvideo.keyframeswithvideo.keyframe_indexes. Editing acceptssource.generation_id,source.url, orsource.data. Extension can identify a prior generation in the start-frame or end-frame field. Looping applies totype: "video"only. The shared aspect-ratio enum has twelve members, but the guide states that no single model-and-type combination accepts all twelve. - Published price table: Luma pricing publishes per-video prices by duration, request type, resolution, and dynamic range. Standard text-to-video or image-to-video costs $0.06 for five seconds and $0.18 for ten seconds at 360p draft; $0.15 and $0.45 at 540p; $0.30 and $0.90 at 720p; and $1.20 and $3.60 at 1080p SDR. Five-second HDR generation costs $0.60 at 720p and $2.40 at 1080p. Five-second HDR plus EXR costs $0.90 at 720p and $3.60 at 1080p. Single-keyframe extension bills one five-second block at $0.15 for 540p, $0.30 for 720p, or $1.20 for 1080p. Reframe is billed per source second at $0.03 for 360p, $0.06 for 540p, $0.12 for 720p, and $0.36 for 1080p. Standard video edits cost $0.54/$1.08 for five/ten seconds at 360p, $0.72/$1.44 at 540p, $1.08/$2.16 at 720p, and $2.16/$4.32 at 1080p. The pricing page also publishes separate HDR and HDR-plus-EXR edit tables. Synchronous request errors are not charged. The page lists refunds for asynchronous
content_moderated,generation_failed, andoutput_not_foundfailures.budget_exhaustedcan receive a partial charge. - Documented job lifecycle and errors: A generation record remains active until completion, failure, or the one-hour job time-to-live. The error guide separates synchronous request rejection from asynchronous failures. Documented asynchronous codes include
content_moderated,generation_failed,output_not_found, andbudget_exhausted. The rate-limit guide defines response headers for current limit, remaining capacity, and reset timing. A completed generation returns a presigned media URL with its own one-hour expiration. - Limits: A submitted video cannot be cancelled. The FAQ states that a job ends on completion, failure, or a one-hour time-to-live. Presigned output URLs expire after one hour. Standard dynamic-range generation has separate five- and ten-second prices. HDR generation is restricted to five seconds and 720p or 1080p. HDR plus EXR has the same duration and resolution restrictions. Single-keyframe extension is standard dynamic range and always bills one five-second block. Reframe is standard dynamic range and bills per second. Input, aspect-ratio, HDR, extension, and output-format combinations are validated by request type. The error guide lists synchronous validation errors and asynchronous failure codes.
- Regions and availability: Pay-as-you-go uses shared capacity. Public documentation does not list a customer-selectable video-processing region for
ray-3.2. - Published support and SLA terms: Pay-as-you-go has no published latency SLA. Luma offers Provisioned Throughput for capacity and latency commitments without contention from shared users. Provisioned image pricing uses requests-per-minute units: one unit equals 1 RPM for
uni-1and 0.4 RPM foruni-1-max.ray-3.2uses a separate video-specific provisioned-capacity class, and its price is available from sales rather than the public price table. The documentation lists request and rate-limit headers for account limits. The public pricing page separates shared pay-as-you-go billing from provisioned capacity and states that video Provisioned Throughput uses sales pricing. - Data retention and training terms: Luma’s April 2026 API terms state that API input and output are not used to train Luma models. The one-hour output-URL expiration is separate from any retention period defined by the contract or account configuration.
- License and usage restrictions: The API terms restrict standalone resale of API access, require disclosure that output is AI-generated, prohibit use of output in AI-training datasets, and require Luma’s prior written consent before publication of performance benchmarks. The terms also contain acceptable-use restrictions.
- Provenance standards: The public Ray 3.2 documentation does not identify a universal signed provenance format for every output. The API returns generation identifiers and request records that identify the Luma generation.
- Primary sources: Ray 3.2 model and control matrix, pay-as-you-go and Provisioned Throughput pricing, error handling, FAQ and output lifetime, rate limits, and Luma API terms.
4. Alibaba Cloud Model Studio: Wan 2.7
- Company and provider: Alibaba Cloud provides Wan models through Model Studio and the DashScope API.
- Current models: International deployment lists
wan2.7-t2v,wan2.7-i2v,wan2.7-r2v, andwan2.7-videoedit. Versioned snapshots includewan2.7-t2v-2026-06-12,wan2.7-t2v-2026-04-25,wan2.7-i2v-2026-04-25, andwan2.7-r2v-2026-06-12. The catalog also lists Wan 2.6, Wan 2.5 preview, Wan 2.2, and Wan 2.1 variants by region and operation. First-frame-only image-to-video models have separatei2v-flash,i2v, and older versioned rows. The catalog separates text-to-video, general image-to-video, first-frame image-to-video, first-and-last-frame generation, reference-to-video, editing, motion transfer, character swap, and digital-human operations. - API endpoints, inputs, and outputs: The Wan video API uses asynchronous task creation and polling. The Singapore endpoint is
https://dashscope-intl.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis; Virginia useshttps://dashscope-us.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis; Beijing useshttps://dashscope.aliyuncs.com/api/v1/services/aigc/video-generation/video-synthesis. Requests requireX-DashScope-Async: enable. The returnedtask_idis valid for 24 hours. Wan 2.7 text-to-video accepts text and supported audio inputs. Image-to-video modes accept text, image, audio, and video inputs according to the operation. Outputs are 720p or 1080p H.264 MP4 at 30 fps for the main Wan 2.7 generation endpoints. - Documented controls: Text-to-video includes prompt, resolution, duration, prompt rewriting, multi-shot settings, and audio inputs where supported. Image-to-video includes first frame, first and last frames, continuation, and continuation with last-frame control. Reference-to-video supports multiple entities and a voice timbre for each entity. General editing covers object and clothing replacement, removal, video extension, frame expansion, and multi-image references. Separate Model Studio operations cover motion transfer, character swapping, image-to-action, style transformation, lip replacement, and digital humans.
- Additional documented specifications: The International Wan 2.7 text-to-video table lists text and audio input, audio-video output, multi-shot narrative, audio-video synchronization, 720p and 1080p, integer duration from 2 to 15 seconds, 30 fps, and H.264 MP4. The Wan 2.7 image-to-video table adds image and video input and lists continuation and last-frame control. The Wan 2.7 reference-to-video table lists text, image, video, and audio input, multi-entity references, entity-specific voice timbre, duration from 2 to 10 seconds, and the same output resolutions and encoding.
- Published price table: Model Studio pricing lists International Wan 2.7 text-to-video and image-to-video at $0.10 per output second for 720p and $0.15/sec for 1080p. Reference-to-video lists $0.10/sec for 720p and $0.15/sec for 1080p under an input-and-output billing label. General video editing uses the same resolution rates and bills both input and output video duration. Failed video-edit requests are not billed and do not consume the listed free quota. The International free quota is 50 output seconds for listed Wan 2.7 generation models and remains valid for 90 days after Model Studio activation. Chinese-mainland Wan 2.7 text-to-video and image-to-video list $0.086012/sec for 720p and $0.143353/sec for 1080p. Global Wan 2.6 in Frankfurt and Virginia lists the same $0.086012 and $0.143353 rates. The US-specific
wan2.6-t2v-usandwan2.6-i2v-usrows list $0.10/sec and $0.15/sec. Internationalwan2.6-i2v-flashlists audio-video output at $0.05/sec for 720p and $0.075/sec for 1080p, and silent output at $0.025/sec and $0.0375/sec. - Documented job lifecycle and errors: HTTP task creation returns a
task_id. Task retrieval uses the asynchronous task endpoint associated with the selected deployment region. The task ID remains valid for 24 hours. The API documentation states that duplicate task creation is separate from polling an existing ID. Successful task data contains the output URL and task metadata. Failed requests do not consume the free quota under the billing rules for the listed video operations. - Limits: International
wan2.7-t2vandwan2.7-i2vlist integer durations from 2 through 15 seconds.wan2.7-r2vlists durations from 2 through 10 seconds. The text-to-video API states that task IDs remain valid for 24 hours and that text-to-video tasks generally take 1 to 5 minutes. Model, API key, and endpoint region must match. HTTP requests requireX-DashScope-Async: enable; synchronous calls return an unsupported-call error. Supported duration, audio behavior, prompt length, and resolution differ by model version and operation. The text-to-video API documents Chinese and English prompts and automatic truncation beyond a model’s character limit. - Regions and availability: Global scope can schedule inference worldwide while storing static data in the selected Virginia or Frankfurt region. International scope can schedule inference worldwide except Chinese mainland and stores static data in Singapore. US scope restricts compute to the United States and stores static data in Virginia. Chinese-mainland scope restricts compute and storage to Beijing. The current tables list Wan 2.7 primarily in International and Chinese-mainland scope; Global and US tables list Wan 2.6 variants for several generation operations.
- Published support and SLA terms: Account quotas and support depend on the Alibaba Cloud account and purchased support plan. The public Wan pages document asynchronous task behavior but do not publish a universal generation-completion SLA for the models listed here.
- Data retention and training terms: Data processing is governed by Alibaba Cloud’s service terms, data-processing agreement, Model Studio terms, and selected deployment scope. The regional documentation distinguishes static-data storage location from the geographic scheduling of inference compute.
- License and usage restrictions: Alibaba Cloud service terms and acceptable-use rules apply. Character replacement, voice, digital-human, and reference inputs are also subject to the customer’s rights and consent obligations under those terms and applicable law.
- Provenance standards: The public Wan overview does not state that all Wan 2.7 outputs carry one common signed provenance manifest. The asynchronous API records the model ID, task ID, deployment scope, and region.
- Primary sources: Model Studio video-generation catalog and regional scopes, Model Studio pricing, and the operation-specific API references linked from the catalog for text-to-video, image-to-video, reference-to-video, and general video editing. Alibaba Cloud service terms and the data-processing agreement are available in the account’s legal-document set.
5. Adobe Firefly Services: Generate Video API
- Company and provider: Adobe provides Generate Video through Firefly Services.
- Current models: Adobe’s Firefly Services changelog records the Generate Video API launch in June 2025. Public API documentation identifies the service as Generate Video but does not publish a dated model ID in the same format used by Google, OpenAI, or AWS.
- API endpoints, inputs, and outputs: Firefly Services uses Adobe Developer Console client credentials and asynchronous jobs. Authentication uses the project’s client ID and access token. The initial generation request returns a job record rather than the finished media, and later job retrieval returns completion data and the output reference. Adobe’s technical usage notes list 16:9 outputs at 960×540, 1280×720, and 1920×1080; 9:16 outputs at 540×960, 720×1280, and 1080×1920; and 1:1 outputs at 540×540, 720×720, and 1080×1080. The API documentation and schema define the available request fields separately from features in the Firefly web and Creative Cloud interfaces.
- Documented controls: Public usage notes document aspect ratio and resolution. The API request schema defines prompt, source-media, duration, and generation fields available to the provisioned account. Adobe does not publish a stable seed or dated model-version field on the general usage-notes page.
- Additional documented specifications: Adobe’s changelog separates Firefly Services API releases from features in Creative Cloud applications. The Generate Video usage page provides exact pixel dimensions instead of only resolution labels. Landscape, portrait, and square each have 540-, 720-, and 1080-class outputs. Firefly Services authentication is managed through an Adobe Developer Console project. The service is part of Adobe’s enterprise API product family, which also includes image generation, image editing, and other creative automation endpoints documented separately.
- Published price table: Adobe does not publish a public dollar-per-second rate for Generate Video. Firefly Services consumes Operations. The Operations rate for each action appears in the customer’s Admin Console or order documents. Adobe’s Shared Credit product description states that one API action can consume more than one Operation. Shared Credits are allocated to an organization and consumed by eligible API activity at product-specific rates. Creative Cloud subscription credits and Firefly web credits are separate from the Firefly Services API rate card. The public product description does not assign one universal Operations cost to every API action.
- Documented job lifecycle and errors: Generate Video follows the Firefly Services asynchronous pattern. The create request produces a job reference, and status retrieval reports the later result. Adobe’s authentication layer can return client-credential and access-token errors before job creation. The documented per-minute and daily request limits can produce rate-limit responses. Content-policy handling and generation failure are represented in the job or API error response according to the endpoint schema.
- Limits: Adobe publishes a default limit of four requests per minute and 9,000 requests per day. At the default start rate, 1,000 submissions require at least 250 minutes. The daily ceiling is separate from the per-minute limit. Higher limits are available through an Adobe account manager. The supported-dimensions table contains nine aspect-ratio and resolution combinations. The general usage-notes page does not publish a batch-submission endpoint that bypasses the per-minute request limit.
- Regions and availability: Firefly Services requires enterprise provisioning through Adobe Developer Console. Public usage notes do not list a customer-selectable processing-region field for Generate Video.
- Published support and SLA terms: Firefly Services uses Adobe enterprise support and account management. The public Generate Video usage notes do not publish a task-completion latency SLA. Higher request limits are arranged through the account manager. Support hours, severity definitions, response targets, and service credits depend on the customer agreement.
- Data retention and training terms: Adobe’s Firefly approach statement states that Adobe does not train Firefly on customer content. Adobe states that Firefly models are trained on licensed material, including Adobe Stock, and public-domain material where copyright has expired.
- License and usage restrictions: Adobe states that it does not claim ownership of customer content. Intellectual-property indemnification is available for eligible enterprise customers, plans, features, surfaces, and export events. The Generative AI Product Specific Terms and Firefly product description define eligibility and exclusions. Adobe’s terms distinguish Adobe Firefly models from partner models and distinguish eligible Firefly output from output created through excluded features or plans. The Firefly product description also defines generative-credit and product-entitlement terms separately from Shared Credits used for enterprise APIs.
- Provenance standards: Adobe supports Content Credentials and is a founding member of the Content Authenticity Initiative. Availability and preservation of Content Credentials depend on the Firefly feature, output path, and later media processing.
- Primary sources: Firefly Services changelog, Generate Video usage notes and limits, Shared Credit product description, Adobe’s Firefly training and customer-content statements, Generative AI Product Specific Terms, and Adobe Firefly product description.
6. fal: multi-model gateway
- Company and provider: fal.ai provides a hosted inference platform and gateway for first-party, partner, and third-party models.
- Current models: Current video listings include Veo 3.1, Wan 2.7, Kling 3, Seedance 2, LTX 2.3, PixVerse V6, Grok Imagine 1.5, HappyHorse 1.0, Happy Oyster, and other model families. Model names and endpoint paths identify the upstream family and operation. Examples include
fal-ai/veo3.1,fal-ai/veo3.1/fast,fal-ai/veo3.1/reference-to-video, and Kling V3 Standard, Pro, and 4K paths. The pricing directory also contains older entries such as Wan 2.5 and Kling 2.5 Turbo Pro alongside newer endpoint-specific model pages. - API endpoints, inputs, and outputs: fal supplies REST, JavaScript, and Python clients with direct execution, queue submission, status, result, webhook, and file-upload operations. Queue submission returns a
request_id; status and result requests address that ID under the selected endpoint. Webhook URLs can be attached to a queue submission. Each model endpoint has its own input and output schema. The Kling V3 Standard text-to-video endpoint accepts either a prompt or amulti_promptshot list. Veo endpoints separate text-to-video, image-to-video, first-and-last-frame, reference-to-video, Fast, and extension operations by path. Successful video jobs return a file object with URL, content type, filename, and size where available. Inputs can be public URLs, data URIs, or files uploaded through fal storage when the endpoint accepts files. - Documented controls: Kling V3 Standard text-to-video lists durations from 3 through 15 seconds and multi-prompt shot durations. Its schema requires either
promptormulti_prompt, but not both. Kling image-to-video endpoints list source image, optional tail image, negative prompt, guidance scale, audio, and model-specific element controls. The Pro image-to-video page lists custom element inputs and native audio. Its output aspect ratio follows the start image. Veo 3.1 endpoints list resolution, audio, aspect ratio, source image, first and last frames, references, and extension according to path. Veo generation supports 720p, 1080p, and 4K endpoint configurations in the current fal documentation. Accepted controls and enum values are defined by the individual endpoint schema. - Additional documented specifications: fal model pages identify endpoint category, partner status, commercial-use status, input schema, output schema, price unit, and example request code. The Kling V3 Standard page identifies text-to-video, image-to-video, and motion-control operations. The Veo 3.1 family has distinct standard and Fast endpoint paths and separate text, image, first-and-last-frame, reference, and extension modes. The client documentation exposes synchronous subscription for waiting clients and queue submission for long-running work, with JavaScript and Python examples.
- Published price table: The fal pricing page lists output units as seconds or completed videos.
fal-ai/veo3.1lists $0.20/sec for silent 720p or 1080p, $0.40/sec with audio, $0.40/sec for silent 4K, and $0.60/sec for 4K with audio.fal-ai/veo3.1/fastlists $0.10/sec silent and $0.15/sec with audio at 720p or 1080p. Kling V3 Pro image-to-video lists $0.112/sec without audio, $0.168/sec with audio, and $0.196/sec with audio and voice control. Kling V3 4K text-to-video lists $0.42/sec with or without ordinary native audio. Other endpoint prices appear on their model pages. - Documented job lifecycle and errors: Queue submission returns a request ID. Queue status reports waiting or processing state and can include endpoint logs. Queue result returns the typed output after completion. Webhooks provide a callback path for long-running requests. fal documents concurrency-limit errors and model-specific validation errors separately from completed model output. A request ID belongs to the endpoint used at submission, so status and result URLs include the same endpoint path.
- Limits: Model-specific limits include duration enums, fixed output sizes, input file types, file-size ceilings, and minimum charges. The Kling V3 Standard text-to-video schema lists durations from 3 to 15 seconds. Its image-to-video type definitions list 5- and 10-second variants for several earlier Kling modes. The Kling V3 Pro page lists videos up to 15 seconds. Veo 3.1 pages list up to eight seconds for one generation and separate extension endpoints. Publicly hosted input URLs must be accessible to fal; hosting services can block cross-site requests or automated fetches. Base64 data URIs increase request size and are documented as less suitable for large files. API keys are server credentials and fal’s documentation states that they must not be exposed in browser or mobile client code.
- Regions and availability: Model and mode availability depends on the endpoint and upstream provider. The Kling V3 4K documentation states that its mode is supported only on the Singapore server. fal’s public common API documentation does not expose one region field shared by all models.
- Published support and SLA terms: fal provides queueing, webhooks, status polling, and enterprise sales and support. Public endpoint pages do not publish one completion-time SLA covering every upstream model. Capacity, concurrency, and support commitments can be defined in enterprise agreements.
- Data retention and training terms: fal’s data-retention guide states that request inputs and outputs are stored by default so they remain retrievable through the platform.
X-Fal-Store-IO: 0prevents platform payload storage. The Platform API can delete request payloads and output CDN files. The header does not delete files uploaded to fal’s CDN during processing. Deletion of request payloads does not delete input CDN files that can be shared by other requests. Uploaded input assets and request-payload records therefore have separate deletion paths. Training terms can also depend on the selected upstream model and any enterprise agreement. - License and usage restrictions: fal’s API supplemental terms state that access materials can change in ways that require client updates. Its general terms disclaim output originality and non-infringement and assign input and end-user responsibilities to the customer under the standard terms. Individual endpoint pages identify partner and commercial-use status where supplied. Upstream-model restrictions can also apply.
- Provenance standards: Provenance is model-specific. fal’s common output file object and request ID identify the gateway request, while any C2PA, watermark, or upstream content credential depends on the selected model and endpoint.
- Primary sources: fal pricing, Veo 3.1 endpoint and configuration rates, Kling V3 Standard API schema, Kling V3 Pro image-to-video pricing, data retention and
X-Fal-Store-IO, API supplemental terms, and general terms.
Figure 2. A direct API connects the application to one model provider. A gateway connects the application to several model providers through an additional API, queue, billing, and storage layer.
7. OpenAI Sora 2: migration-only API
- Company and provider: OpenAI provides Sora 2 through the OpenAI Videos API.
- Current models:
sora-2and Sora 2 Pro are deprecated. The Sora 2 model page lists thesora-2alias and dated snapshots includingsora-2-2025-12-08andsora-2-2025-10-06. The catalog describes Sora 2 as video generation with synchronized audio. OpenAI’s discontinuation notice states that the API will end September 24, 2026. - API endpoints, inputs, and outputs: The model page lists
/v1/videos. Sora 2 accepts text and image input and returns video with synchronized audio. Published output sizes include 720×1280 portrait and 1280×720 landscape. The model page marks text and image as input-only modalities and audio and video as output-only modalities. The video endpoint is separate from Responses, Chat Completions, image generation, and audio endpoints. - Documented controls: The Videos API schema defines prompt, input image, duration, output size, model alias or snapshot, and operation-specific video fields. Snapshot IDs provide a dated model reference while available.
- Additional documented specifications: The Sora 2 catalog page marks the model as Legacy and Deprecated. It lists text and image as input-only modalities and audio and video as output-only modalities. The catalog provides pricing, snapshot aliases, endpoint support, and tier-based rate limits on the same model record. Sora 2 Pro appears in the quick comparison as a separate model and price tier. The discontinuation notice separates the April 26, 2026 consumer-app closure from the September 24, 2026 API closure.
- Published price table: The Sora 2 model page lists $0.10 per generated second at 720×1280 or 1280×720. Its quick comparison lists Sora 2 Pro at $0.30 per generated second.
- Documented job lifecycle and errors: The Videos API creates and retrieves video-generation jobs. Request authentication, parameter validation, content-policy handling, rate limits, and processing failures are returned through the API’s documented error format. The deprecation status is independent of individual job state: an otherwise valid job can run while the service remains available, and API access ends on the announced service date.
- Limits: The model page lists rate limits by usage tier: 25 requests per minute at Tier 1, 50 at Tier 2, 125 at Tier 3, 200 at Tier 4, and 375 at Tier 5. Free-tier access is unsupported. Rate limits are account-tier limits and remain subordinate to the September 24, 2026 service end date.
- Regions and availability: The consumer Sora web and app experiences ended April 26, 2026. The Sora API remains available only until September 24, 2026 under the published discontinuation schedule. The model catalog marks Sora 2 and Sora 2 Pro deprecated.
- Published support and SLA terms: OpenAI provides support and a service-status site for API services. The Sora discontinuation notice does not publish a completion-time SLA or service extension beyond September 24, 2026.
- Data retention and training terms: The discontinuation notice states that Sora-associated data will be permanently deleted after discontinuation and after any final export window. The notice describes export access through
sora.chatgpt.com/sunsetfor Sora content and states that the export is delivered after an email notification. The notice also states that a final export window might be offered and that users would receive an email before such a window. - License and usage restrictions: OpenAI’s service terms and usage policies apply while the API remains available. Deprecation does not alter rights or obligations for content generated before shutdown.
- Provenance standards: Sora request records include model alias or snapshot, request ID, and output metadata. The cited model page does not identify a universal C2PA requirement for every Sora 2 API output.
- Primary sources: Sora discontinuation notice, Sora 2 model, price, snapshots, and rate limits, and OpenAI model catalog.
8. Amazon Nova Reel: 1.0 retiring, 1.1 region-limited
- Company and provider: Amazon Web Services provides Nova Reel through Amazon Bedrock.
- Current models: Nova Reel 1.0 uses
amazon.nova-reel-v1:0and is marked legacy with a September 30, 2026 end-of-life date. Nova Reel 1.1 usesamazon.nova-reel-v1:1. - API endpoints, inputs, and outputs: Nova Reel uses Amazon Bedrock
StartAsyncInvoke; the synchronousInvokeModelAPI is unsupported. Version 1.1 accepts text or text plus one initial image and returns video through an S3 destination in the customer account. Output is 1280×720 at 24 fps. Duration uses six-second increments up to two minutes. Longer jobs support automated multi-shot generation from one prompt or manual storyboards with a prompt for each six-second shot. The S3 invocation folder is named for the invocation ID and containsoutput.mp4,manifest.json,generation-status.json, and constituent shots for longer generations. Bedrock status retrieval usesGetAsyncInvoke, and invocation enumeration usesListAsyncInvokes. - Documented controls: Request fields include task type, prompt, optional image, output duration, fixed dimensions, fixed frame rate, output S3 URI, and seed. The seed range is 0 through 2,147,483,646 and defaults to 42. Ordinary text-to-video prompts are limited to 512 characters. Automated multi-shot prompts can contain up to 4,000 characters. Manual storyboards accept up to 512 characters per shot. AWS prompt documentation lists subject, action, environment, lighting, style, and camera motion as prompt components. Manual long-video generation assigns one prompt to each six-second interval.
- Additional documented specifications: The Nova Reel service card defines an “effective” output as one that contains requested prompt content, makes reasonable assumptions for unspecified details, avoids listed composition defects, and meets the evaluator’s safety and fairness standards. AWS states that evaluation uses public benchmarks and proprietary datasets but does not publish a customer approval percentage in the cited card. For long generation, Bedrock stores both the stitched output and individual six-second shots. The model documentation lists text-to-video, text-and-image-to-video, and Content Credentials as supported features.
- Published price table: Amazon Bedrock pricing lists Nova Reel generation at $0.08 per output second for 720p, 24 fps video. Model charges are separate from S3 storage and data-transfer charges.
- Documented job lifecycle and errors:
StartAsyncInvokereturns an invocation ARN.GetAsyncInvokereports status, andListAsyncInvokesenumerates account invocations. AWS documents completed and failed asynchronous states and writes generation status into the selected S3 prefix. The output prefix is associated with one invocation ID. Bedrock authorization errors, invalid model inputs, quota errors, S3 permission failures, and model-generation failures occur at different stages of the request. - Limits: Nova Reel 1.1 supports six-second duration increments and a maximum of two minutes. Fine-tuning and Provisioned Throughput are unsupported. AWS’s quota reference lists three concurrent on-demand requests for Nova Reel 1.1 by default and ten for Nova Reel 1.0. The quota reference marks the listed concurrency quotas as not adjustable. The model supports English prompts. Output resolution and frame rate are fixed at 1280×720 and 24 fps. Video longer than six seconds requires the
amazon.nova-reel-v1:1model ID. - Regions and availability: Nova Reel 1.1 is available only in
us-east-1(Northern Virginia). Nova Reel 1.0 is listed in Northern Virginia, Ireland, and Tokyo until its end-of-life date. The model card lists in-region access and no geo or global inference ID for 1.0. - Published support and SLA terms: AWS documentation states that a six-second generation typically takes about 90 seconds and a two-minute generation about 14 to 17 minutes. These are documented typical times rather than an SLA. Amazon Bedrock uses AWS Support plans and service terms. The Nova Reel model page lists Standard service tier support and no Priority, Flex, or Reserved tier for version 1.0. Generation requires
bedrock:InvokeModeland S3PutObject; AWS also listsbedrock:GetAsyncInvokeandbedrock:ListAsyncInvokesfor status tracking. - Data retention and training terms: The Nova Reel service card states that Amazon Bedrock does not store or review customer prompts or video outputs, does not share them between customers or with third-party model providers, and does not use them to train Amazon Bedrock models, including Nova Reel. Output is written to the customer’s S3 bucket and follows the bucket’s retention, encryption, logging, and access configuration. Bedrock writes the output on behalf of the caller through the IAM and S3 permissions attached to the invocation.
- License and usage restrictions: AWS service terms, the Nova Reel end-user license terms, and the AWS acceptable-use policy apply. Rights to text, images, people, products, trademarks, and other submitted material remain governed by the customer’s licenses and applicable law.
- Provenance standards: Nova Reel 1.1 includes Content Credentials. AWS documentation identifies Content Credentials Verify as a public verification method when the metadata remains present.
- Primary sources: Nova Reel 1.0 model card and end-of-life date, Nova Reel 1.1 specifications, access, typical latency, S3 output, and request fields, AWS quota reference, Nova Reel service card, and Amazon Bedrock pricing.