← All models
Wan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body movements, and professional camera work for film and television applications
Example
Details
- Accepts
- image + audio
Pricing
- Price
- 25 cr / second
Prices in credits (1 credit = $0.01).
Data schema
Input
| Field | Type | Description |
|---|---|---|
| seed | — | Random seed for reproducibility. If None, a random seed is chosen. |
| shift | number | Shift value for the video. Must be between 1.0 and 10.0. |
| promptrequired | string | The text prompt used for video generation. |
| audio_urlrequired | string | The URL of the audio file. |
| image_urlrequired | string | URL of the input image. If the input image does not match the chosen aspect ratio, it is resized and center cropped. |
| num_frames | integer | Number of frames to generate. Must be between 40 to 120, (must be multiple of 4). |
| resolution | stringenum: 480p, 580p, 720p | Resolution of the generated video (480p, 580p, or 720p). |
| video_quality | stringenum: low, medium, high, maximum | The quality of the output video. Higher quality means better visual quality but larger file size. |
| guidance_scale | number | Classifier-free guidance scale. Higher values give better adherence to the prompt but may decrease quality. |
| negative_prompt | string | Negative prompt for video generation. |
| video_write_mode | stringenum: fast, balanced, small | The write mode of the output video. Faster write mode means faster results but larger file size, balanced write mode is a good compromise between speed and quality, and small write mode is the slowest but produces the smallest file size. |
| frames_per_second | — | Frames per second of the generated video. Must be between 4 to 60. When using interpolation and `adjust_fps_for_interpolation` is set to true (default true,) the final FPS will be multiplied by the number of interpolated frames plus one. For example, if the generated frames per second is 16 and the … |
| num_inference_steps | integer | Number of inference steps for sampling. Higher values give better quality but take longer. |
| enable_safety_checker | boolean | If set to true, input data will be checked for safety before processing. Disabling it requires account authorization; unauthorized requests are always checked. |
| enable_output_safety_checker | boolean | If set to true, output video will be checked for safety after generation. |
Output
| Field | Type | Description |
|---|---|---|
| video | — | The generated video file. |