← All models
Seedance 2.0 Fast Reference to Video
seedance-2.0-fast-reference-to-video
Video generationby ByteDance
ByteDance's most advanced reference-to-video model, fast tier. Lower latency and cost with up to 9 images, 3 videos, and 3 audio clips as inputs.
Example
Details
- Accepts
- text + image + video + audio
Pricing
- Input
- 1400 cr / 1M tokens
Prices in credits (1 credit = $0.01).
Data schema
Input
| Field | Type | Description |
|---|---|---|
| codec | stringenum: auto, H264, H265 | 'auto' retains default video codec behaviour; 'H264' uses H.264; 'H265' uses H.265. |
| promptrequired | string | The text prompt describing the desired motion and action for the video. |
| duration | stringenum: auto, 4, 5, 6, 7, 8… | Duration of the video in seconds. Supports 4 to 15 seconds, or auto to let the model decide based on the prompt. |
| image_url | string | The URL of the starting frame image to animate. Supported formats: JPEG, PNG, WebP. Max 30 MB. |
| resolution | stringenum: 480p, 720p | Video resolution - 480p for faster generation, 720p for balance. |
| end_user_id | — | The unique user ID of the end user. |
| aspect_ratio | stringenum: auto, 21:9, 16:9, 4:3, 1:1, 3:4… | The aspect ratio of the generated video. Use 16:9 for landscape, 9:16 for portrait/vertical, 1:1 for square, 21:9 for ultrawide cinematic, or auto to infer from the input image. |
| bitrate_mode | stringenum: standard, high | Output bitrate mode. 'high' requests a higher-quality, larger-file encode from the model; 'standard' uses the default bitrate. |
| end_image_url | — | The URL of the image to use as the last frame of the video. When provided, the generated video will transition from the starting image to this ending image. Supported formats: JPEG, PNG, WebP. Max 30 MB. |
| generate_audio | boolean | Whether to generate synchronized audio for the video, including sound effects, ambient sounds, and lip-synced speech. The cost of video generation is the same regardless of whether audio is generated or not. |
| audio_urls | array | Reference audio to guide video generation. Refer to them in the prompt as @Audio1, @Audio2, etc. Supported formats: MP3, WAV. Up to 3 files, combined duration must not exceed 15 seconds. Max 15 MB per file. At least one reference image or video is required. |
| image_urls | array | Reference images to guide video generation. Refer to them in the prompt as @Image1, @Image2, etc. Supported formats: JPEG, PNG, WebP. Max 30 MB per image. Up to 9 images. Total files across all modalities must not exceed 12. |
| video_urls | array | Reference videos to guide video generation. Refer to them in the prompt as @Video1, @Video2, etc. Supported formats: MP4, MOV. Up to 3 videos, combined duration must be between 2 and 15 seconds, total size under 50 MB. Each video must be between ~480p (640x640) and ~720p (834x1112) in resolution. |
| seed | integer | Random seed. Set for reproducible generation. |
| image | string | Input image for image-to-video generation (first frame). Cannot be combined with reference images. |
| last_frame_image | string | Input image for last frame generation. Only works if a first frame image is also provided. Cannot be combined with reference images. |
| reference_audios | array | Reference audio files (up to 3, total duration max 15s) for audio-driven generation and lip-sync. Requires at least one reference image or video. Reference them in your prompt as [Audio1], [Audio2], etc. |
| reference_images | array | |
| reference_videos | array | Reference videos (up to 3, total duration max 15s) for motion transfer, style reference, and editing. Reference them in your prompt as [Video1], [Video2], etc. |
Output
| Field | Type | Description |
|---|---|---|
| seed | integer | The seed used for generation. |
| video | — | The generated video file. |