infery
← All models

Kling Video v3 Text to Video [Standard]

kling-video-v3-standard-text-to-video

Video generationby Kling

Kling 3.0 Standard: Top-tier text-to-video with cinematic visuals, fluid motion, and native audio generation, with multi-shot support.

Example

Details

Accepts
text + image

Pricing

Price
17.5 cr / second

Prices in credits (1 credit = $0.01).

Data schema

Input

FieldTypeDescription
promptText prompt for video generation. Either prompt or multi_prompt must be provided, but not both.
durationstringenum: 3, 4, 5, 6, 7, 8The duration of the generated video in seconds
elementsElements (characters/objects) to include in the video. Each example can either be an image set (frontal + reference images) or a video. Reference in prompt as @Element1, @Element2, etc.
cfg_scalenumber The CFG (Classifier Free Guidance) scale is a measure of how close you want the model to stick to your prompt.
shot_typestringenum: customize, intelligentThe type of multi-shot video generation. 'intelligent' lets the model automatically determine shot structure.
multi_promptList of prompts for multi-shot video generation. If provided, divides the video into multiple shots.
end_image_urlURL of the image to be used for the end of the video
generate_audiobooleanWhether to generate native audio for the video. Supports Chinese and English voice output. Other languages are automatically translated to English. For English speech, use lowercase letters; for acronyms or proper nouns, use uppercase.
negative_promptstring
start_image_urlstringURL of the image to be used for the video
aspect_ratiostringenum: 16:9, 9:16, 1:1The aspect ratio of the generated video frame

Output

FieldTypeDescription
videoThe generated video