Video models
410 models. Compare AI video models by what they take as input — a text prompt, a source image, or an existing clip to edit or upscale — and by clip length and resolution.
Price ranks a model against others of the same modality — a second of video generation and a second of image generation aren't the same unit of work, so thirds are computed within each modality, not across all of them.
410 of 410 models

AI Avatar Multi
AI Avatar · Video edit
MultiTalk model generates a multi-person conversation video from an image and audio files.
25 cr / second
Try in Studio →
AI Avatar Multi Text
AI Avatar · Video edit
MultiTalk model generates a multi-person conversation video from an image and text inputs.
25 cr / second
Try in Studio →
AI Avatar Single Text
AI Avatar · Video edit
MultiTalk model generates a talking avatar video from an image and text.
25 cr / second
Try in Studio →
Avatars
Veed · Video generation
Generate high-quality videos with UGC-like avatars from audio
0.072 cr / second
Try in Studio →
Avatars Text to Video
Argil · Video generation
High-quality avatar videos that feel real, generated from your text
0.156 cr / second
Try in Studio →
Avatar X
Mirage Api · Video generation
The Avatar X API offers access to Mirage's most advanced generation model yet, delivering industry-leading identity preservation and…
37.5 cr / second
Try in Studio →
Bernini-R Edit Video
Bernini R · Video generation
Edit any video with a natural-language instruction using Bernini-R, changing objects, weather, background, or camera angle while keeping…
0.156 cr / second
Try in Studio →
Bernini-R Reference Edit Video
Bernini R · Video edit
Edit a video guided by reference images with Bernini-R, bringing an object, material, background, style, or weather from a reference image…
0.156 cr / second
Try in Studio →
Bernini-R Reference to Video
Bernini R · Video edit
Turn up to five reference images into one continuous, consistent video with Bernini-R, with smooth, stable camera motion and no scene cuts.
0.156 cr / second
Try in Studio →
Bernini-R Text to Video
Bernini R · Video generation
Generate high-quality video from a text prompt with Bernini-R, ByteDance's unified video generation and editing model.
0.156 cr / second
Try in Studio →
Birefnet
Birefnet · Video generation
Video background removal version of bilateral reference framework (BiRefNet) for high-resolution dichotomous image segmentation (DIS)
0.100 cr / second
Try in Studio →
Bria's VRMBG 3.0
Bria · Video generation
Remove backgrounds from any video with Bria's VRMBG 3.0.
0.00375 cr / second
Try in Studio →
Bria's VRMBG 3.0 Realtime
Bria · Video generation
Remove video backgrounds in real time with Bria’s VRMBG 3.0 model.
0.156 cr / second
Try in Studio →
Bria Video Eraser
Bria · Video generation
A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and…
17.5 cr / second
Try in Studio →
Bria Video Eraser
Bria · Video generation
A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and…
17.5 cr / second
Try in Studio →
Bria Video Eraser Erase Mask
Bria · Video generation
A high-fidelity capability for erasing unwanted objects, people, or visual elements from videos while maintaining aesthetic quality and…
17.5 cr / second
Try in Studio →
Bytedance Dreamactor V2
ByteDance · Video edit
Transfer motion from a video to characters in an image using Dreamactor v2. Great performance for non-human and multiple characters
6.25 cr / second
Try in Studio →
Bytedance Omnihuman V1.5
ByteDance · Video edit
Omnihuman v1.5 is a new and improved version of Omnihuman. It generates video using an image of a human figure paired with an audio file.
20 cr / second
Try in Studio →
Bytedance Seedance V1.5 Pro Text To Video
ByteDance · Video generation
Generate videos with audio with Seedance 1.5
150 cr / 1M tokens
Try in Studio →
Bytedance Seedance V1 Pro Fast Text To Video
ByteDance · Video generation
Text to Video endpoint for Seedance 1.0 Pro Fast, a next-generation video model designed to deliver maximum performance at minimal cost
125 cr / 1M tokens
Try in Studio →
Bytedance Upscaler Upscale Video
Bytedance Upscaler · Video upscale
Upscale videos with Bytedance's video upscaler.
0.900 cr / second
Try in Studio →
Cosmos 3 Super Image to Video
NVIDIA · Video edit
Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from…
6.25 cr / second
Try in Studio →
Cosmos Predict 2.5 2B
NVIDIA · Video generation
Generate video from text and videos using NVIDIA's 2B Cosmos Post-Trained Model
25 cr / video
Try in Studio →
Cosmos Predict 2.5 2B Distilled
NVIDIA · Video generation
Generate video from text and videos using NVIDIA's 2B Cosmos Distilled Model
10 cr / video
Try in Studio →
Davinci Magihuman
Davinci Magihuman · Video edit
Expressive facial performance, natural speech-expression coordination, realistic body motion, and accurate audio-video synchronization with…
6.25 cr / second
Try in Studio →
Depth Anything Video
Depth Anything Video · Video generation
Generates depth maps from video using Video Depth Anything (CVPR 2025).
5 cr / second
Try in Studio →
DWPose Pose Prediction
Dwpose · Video generation
Predict poses from videos.
0.075 cr / second
Try in Studio →
EchoMimic V3
Echomimic V3 · Video edit
EchoMimic V3 generates a talking avatar model from a picture, audio and text prompt.
25 cr / second
Try in Studio →
Editto
Editto · Video generation
Edit videos using instruction-based prompting using Editto model!
10 cr / second
Try in Studio →
ElevenLabs Dubbing
Elevenlabs · Video generation
Generate dubbed videos or audios using ElevenLabs Dubbing feature!
75 cr / minute
Try in Studio →
Fabric 1.0
Veed · Video edit
VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video
0.021 cr / second
Try in Studio →
Fabric 1.0 Fast
Veed · Video edit
VEED Fabric 1.0 is an image-to-video API that turns any image into a talking video
0.021 cr / second
Try in Studio →
Ffmpeg Api
infery.ai · Video generation
Use ffmpeg capabilities to merge 2 or more videos.
0.021 cr / second
Try in Studio →
FFmpeg API Compose
infery.ai · Video generation
Compose videos from multiple media sources using FFmpeg API.
0.025 cr / second
Try in Studio →
Ffmpeg Api Merge Audio-Video
infery.ai · Video generation
Merge videos with standalone audio files or audio from video files.
0.025 cr / second
Try in Studio →
FILM
Film · Video generation
Interpolate videos with FILM - Frame Interpolation for Large Motion
0.163 cr / second
Try in Studio →
Flashhead
Flashhead · Video edit
SoulX-FlashHead is a unified 1.3B-parameter framework designed for high-fidelity, infinite-length, and real-time streaming portrait video…
0.625 cr / second
Try in Studio →
Flashtalk
Flashtalk · Video edit
Audio-driven talking avatar generation powered by the SoulX-FlashTalk 14B model.
2.5 cr / second
Try in Studio →
Flashvsr
Flashvsr · Video upscale
Upscale your videos using FlashVSR with the fastest speeds!
0.063 cr / megapixel
Try in Studio →
Flux 3 Draft Enhance
Blackforestlabs · Video generation
FLUX.3 is Black Forest Labs' frontier audio/video model.
10.63 cr / second
Try in Studio →
Flux 3 Extend Video Draft
Blackforestlabs · Video generation
FLUX.3 is Black Forest Labs' frontier audio/video model.
7.5 cr / second
Try in Studio →
Flux 3 First Last Frame to Video
Blackforestlabs · Video edit
FLUX.3 is Black Forest Labs' frontier video model.
10.63 cr / second
Try in Studio →
Flux 3 First Last Frame to Video Draft
Blackforestlabs · Video edit
FLUX.3 is Black Forest Labs' frontier audio/video model.
3.75 cr / second
Try in Studio →
Flux 3 Image to Video
Blackforestlabs · Video edit
FLUX.3 is Black Forest Labs' frontier video model.
10.63 cr / second
Try in Studio →
Flux 3 Image To Video Draft
Blackforestlabs · Video edit
FLUX.3 is Black Forest Labs' frontier audio/video model.
3.75 cr / second
Try in Studio →
Flux 3 Keyframes To Video Draft
Blackforestlabs · Video edit
FLUX.3 is Black Forest Labs' frontier audio/video model.
3.75 cr / second
Try in Studio →
Flux 3 Text to Video
Blackforestlabs · Video generation
FLUX.3 is Black Forest Labs' frontier video model.
10.63 cr / second
Try in Studio →
Flux 3 Text To Video Draft
Blackforestlabs · Video generation
FLUX.3 is Black Forest Labs' frontier audio/video model.
3.75 cr / second
Try in Studio →
Framepack
Framepack · Video edit
Framepack is an efficient Image-to-video model that autoregressively generates videos.
4.16 cr / second
Try in Studio →
Framepack
Framepack · Video edit
Framepack is an efficient Image-to-video model that autoregressively generates videos.
4.16 cr / second
Try in Studio →
Framepack F1
Framepack · Video edit
Framepack is an efficient Image-to-video model that autoregressively generates videos.
4.16 cr / second
Try in Studio →Gen 4.5
Runwayml · Video generation
Runway Gen-4.5 is Runway's latest text-to-video model with high motion quality, prompt adherence, and visual fidelity.
15 cr / second
Try in Studio →
Grok Imagine Reference to Video
xAI · Video edit
Generate videos using multiple reference images with xAI's Grok Imagine video model
6.25 cr / second
Try in Studio →
Grok Imagine Video
xAI · Video generation
Edit videos using xAI's Grok Imagine
6.25 cr / second
Try in Studio →
Grok Imagine Video
xAI · Video generation
Grok Imagine Video is xAI's image-to-video model that animates still images into short videos with synchronized audio.
6.25 cr / second
Try in Studio →Grok Imagine Video 1.5
xAI · Video generation
xAI's Grok Imagine Video 1.5 (preview) animates still images into short videos with synchronized audio.
10 cr / second
Try in Studio →Video, on Infery.ai
Video models split into three catalogue categories: video generation (a prompt or a still image becomes a new clip), video edit (an existing clip is modified — extended, restyled, or given a different ending), and video upscale (an existing clip is resampled to a higher resolution without regenerating its content). All three run through the same OpenAI-compatible endpoint and the same credit balance.
It is a catalogue, not one model
The catalogue splits mainly on primary input: some models take only a text prompt, others require a source image or an existing clip as the starting point. Clip length and output resolution vary by model and are listed on each model's detail page next to its price — a per-second rate makes a longer clip cost proportionally more, so price and duration have to be read together, not the rate alone.
Open the Studio with a video model already picked
Free trial credits, no card. Filter above, or start from a blank chat and pick as you go.



