infery
← All models

Seedance 2.0 Fast Reference to Video

seedance-2.0-fast-reference-to-video

Video generationby ByteDance

ByteDance's most advanced reference-to-video model, fast tier. Lower latency and cost with up to 9 images, 3 videos, and 3 audio clips as inputs.

Example

Details

Accepts
text + image + video + audio

Pricing

Input
1400 cr / 1M tokens

Prices in credits (1 credit = $0.01).

Data schema

Input

FieldTypeDescription
codecstringenum: auto, H264, H265'auto' retains default video codec behaviour; 'H264' uses H.264; 'H265' uses H.265.
promptrequiredstringThe text prompt describing the desired motion and action for the video.
durationstringenum: auto, 4, 5, 6, 7, 8…Duration of the video in seconds. Supports 4 to 15 seconds, or auto to let the model decide based on the prompt.
image_urlstringThe URL of the starting frame image to animate. Supported formats: JPEG, PNG, WebP. Max 30 MB.
resolutionstringenum: 480p, 720pVideo resolution - 480p for faster generation, 720p for balance.
end_user_id—The unique user ID of the end user.
aspect_ratiostringenum: auto, 21:9, 16:9, 4:3, 1:1, 3:4…The aspect ratio of the generated video. Use 16:9 for landscape, 9:16 for portrait/vertical, 1:1 for square, 21:9 for ultrawide cinematic, or auto to infer from the input image.
bitrate_modestringenum: standard, highOutput bitrate mode. 'high' requests a higher-quality, larger-file encode from the model; 'standard' uses the default bitrate.
end_image_url—The URL of the image to use as the last frame of the video. When provided, the generated video will transition from the starting image to this ending image. Supported formats: JPEG, PNG, WebP. Max 30 MB.
generate_audiobooleanWhether to generate synchronized audio for the video, including sound effects, ambient sounds, and lip-synced speech. The cost of video generation is the same regardless of whether audio is generated or not.
audio_urlsarrayReference audio to guide video generation. Refer to them in the prompt as @Audio1, @Audio2, etc. Supported formats: MP3, WAV. Up to 3 files, combined duration must not exceed 15 seconds. Max 15 MB per file. At least one reference image or video is required.
image_urlsarrayReference images to guide video generation. Refer to them in the prompt as @Image1, @Image2, etc. Supported formats: JPEG, PNG, WebP. Max 30 MB per image. Up to 9 images. Total files across all modalities must not exceed 12.
video_urlsarrayReference videos to guide video generation. Refer to them in the prompt as @Video1, @Video2, etc. Supported formats: MP4, MOV. Up to 3 videos, combined duration must be between 2 and 15 seconds, total size under 50 MB. Each video must be between ~480p (640x640) and ~720p (834x1112) in resolution.
seedintegerRandom seed. Set for reproducible generation.
imagestringInput image for image-to-video generation (first frame). Cannot be combined with reference images.
last_frame_imagestringInput image for last frame generation. Only works if a first frame image is also provided. Cannot be combined with reference images.
reference_audiosarrayReference audio files (up to 3, total duration max 15s) for audio-driven generation and lip-sync. Requires at least one reference image or video. Reference them in your prompt as [Audio1], [Audio2], etc.
reference_imagesarray
reference_videosarrayReference videos (up to 3, total duration max 15s) for motion transfer, style reference, and editing. Reference them in your prompt as [Video1], [Video2], etc.

Output

FieldTypeDescription
seedintegerThe seed used for generation.
video—The generated video file.