Alibaba models
155 models on Infery.ai.
CosyVoice V2
cosyvoice-v2
CosyVoice V3 Flash
cosyvoice-v3-flash
CosyVoice V3 Plus
cosyvoice-v3-plus

Happy Horse
happy-horse-text-to-video
Generate 1080p video with synchronized native audio from a text prompt. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4. Duration: 3–15s.

Happy Horse
happy-horse-reference-to-video
Generate 1080p video with synchronized native audio from a text prompt and references. Aspect ratios: 16:9, 9:16, 1:1, 4:3, 3:4.
Happyhorse 1.0
happyhorse-1.0
Generate video from text or animate an image with Happy Horse 1.0 by Alibaba. 720p/1080p, 3-15s, five aspect ratios.
Happyhorse 1.1
happyhorse-1.1
Generate video from text, animate an image, or combine reference images with Happy Horse 1.1 by Alibaba.
HappyHorse 1.1 Image-to-Video
happyhorse-1-1-i2v

Happy Horse 1.1 Reference to Video
happy-horse-v1.1-reference-to-video
Happy Horse 1.1 is Alibaba's #1-ranked video model.
HappyHorse 1.1 Ref-to-Video
happyhorse-1-1-r2v

Happy Horse 1.1 Text to Video
happy-horse-v1.1-text-to-video
Happy Horse 1.1 is Alibaba's #1-ranked video model.
HappyHorse 1.1 Text-to-Video
happyhorse-1-1-t2v
Paraformer V2
paraformer-v2
Qwen3.5 Flash
qwen3.5-flash
1M ctx
Qwen3.5 Omni Flash
qwen3.5-omni-flash
262K ctx
Qwen3.5 Omni Plus
qwen3.5-omni-plus
262K ctx
Qwen3.5 Plus
qwen3.5-plus
1M ctx
Qwen3.6 Flash
qwen3.6-flash
Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series.
1M ctx
Qwen3.7 Max
qwen3.7-max
Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. 1,000,000 token context window, maximum output of 65,536 tokens.
1M ctx
Qwen3.7 Plus
qwen3.7-plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. 1,000,000 token context window, maximum output of 65,536 tokens.
1M ctx
Qwen3 ASR Flash
qwen3-asr-flash
Qwen3 Coder Flash
qwen3-coder-flash
Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus.
1M ctx
Qwen3 Coder Plus
qwen3-coder-plus
Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B.
1M ctx
Qwen3 Max
qwen3-max
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual…
262K ctx
Qwen3 Omni Flash
qwen3-omni-flash
66K ctx
Qwen3 Rerank
qwen3-rerank
33K ctx

Qwen 3 TTS - Clone Voice [0.6B]
qwen-3-tts-clone-voice-0.6b
Clone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create…

Qwen 3 TTS - Clone Voice [1.7B]
qwen-3-tts-clone-voice-1.7b
Clone your voices using Qwen3-TTS Clone-Voice model with zero shot cloning capabilities and use it on text-to-speech models to create…

Qwen 3 TTS - Text to Speech [0.6B]
qwen-3-tts-text-to-speech-0.6b
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice…

Qwen 3 TTS - Text to Speech [1.7B]
qwen-3-tts-text-to-speech-1.7b
Bring speech to your texts using Qwen3-TTS Custom-Voice model with pre-trained voices or use your custom voice with Qwen3-TTS Clone Voice…

Qwen 3 TTS - Voice Design [1.7B]
qwen-3-tts-voice-design-1.7b
Create custom voices using Qwen3-TTS Voice Design model and later use Clone Voice model to create your own voices!
Qwen3 VL Flash
qwen3-vl-flash
262K ctx
Qwen3 VL Plus
qwen3-vl-plus
262K ctx

Qwen Audio 3.0 TTS (Flash)
qwen-audio-3-tts
Generate natural multilingual speech from text with fast voice and language control using Qwen Audio 3.0 TTS Flash.
Qwen Flash
qwen-flash
1M ctx

Qwen Image
qwen-image
Qwen Image 2.0
qwen-image-2.0
Qwen Image 2.0 Pro
qwen-image-2.0-pro

Qwen Image 2512
qwen-image-2512-lora
LoRA inference endpoint for Qwen Image 2512, an improved version of Qwen Image with better text rendering, finer natural textures, and more…

Qwen Image 2512
qwen-image-2512
Qwen Image 2512 is an improved version of Qwen Image with better text rendering, finer natural textures, and more realistic human…

Qwen Image Edit 2509
qwen-image-edit-2509
Endpoint for Qwen's Image Editing Plus model also known as Qwen-Image-Edit-2509.

Qwen Image Edit 2509 Lora
qwen-image-edit-2509-lora
LoRA endpoint for the Qwen Image Edit 2509 model.

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-remove-lighting
Remove existing lighting and apply soft, even illumination

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-add-background
Add a realistic scene behind the object with white background

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-shirt-design
Apply designs/graphics onto people's shirts

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-integrate-product
Blend products into backgrounds with automatic perspective and lighting correction

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-remove-element
Remove unwanted elements (objects, people, text) while maintaining image consistency

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-lighting-restoration
Removes harsh shadows and light spots from images, replacing them with soft, even, natural-looking illumination.

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-group-photo
Create group photos

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-next-scene
Create cinematic transitions and scene progressions (camera movements, framing changes)

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-face-to-full-portrait
Generate full portrait from a cropped face photo

Qwen Image Edit 2509 Lora Gallery
qwen-image-edit-2509-lora-gallery-multiple-angles
Precise camera position and angle control (rotation, zoom, vertical movement)

Qwen Image Edit 2511
qwen-image-edit-2511-lora
Endpoint for Qwen's Image Editing 2511 model with LoRa support.

Qwen Image Edit 2511
qwen-image-edit-2511
Endpoint for Qwen's Image Editing 2511 model.

Qwen Image Edit 2511 Multiple Angles
qwen-image-edit-2511-multiple-angles
Generates same scene from different angles (azimuth/elevation) with Qwen image Edit 2511 and the Lora Multiple Angles

Qwen Image Edit Lora
qwen-image-edit-lora
LoRA inference endpoint for the Qwen Image Editing model.
Qwen Image Edit Max
qwen-image-edit-max

Qwen Image Edit Plus
qwen-image-edit-plus

Qwen Image Edit Plus Lora
qwen-image-edit-plus-lora
LoRA endpoint for the Qwen Image Edit Plus model.

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-group-photo
Create group photos

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-next-scene
Create cinematic transitions and scene progressions (camera movements, framing changes)

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-integrate-product
Blend products into backgrounds with automatic perspective and lighting correction

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-add-background
Add a realistic scene behind the object with white background

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-face-to-full-portrait
Generate full portrait from a cropped face photo

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-remove-lighting
Remove existing lighting and apply soft, even illumination

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-remove-element
Remove unwanted elements (objects, people, text) while maintaining image consistency

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-lighting-restoration
Removes harsh shadows and light spots from images, replacing them with soft, even, natural-looking illumination.

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-shirt-design
Apply designs/graphics onto people's shirts

Qwen Image Edit Plus Lora Gallery
qwen-image-edit-plus-lora-gallery-multiple-angles
Precise camera position and angle control (rotation, zoom, vertical movement)

Qwen Image Layered
qwen-image-layered-lora
Qwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers. Use loras to get your custom outputs.

Qwen Image Layered
qwen-image-layered
Qwen-Image-Layered is a model capable of decomposing an image into multiple RGBA layers.

Qwen Image Max
qwen-image-max
Qwen Image Plus
qwen-image-plus
Qwen Long
qwen-long
10M ctx
Qwen Plus
qwen-plus
Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.
1M ctx
Qwen Text Embedding v3
qwen-text-embedding-v3
8K ctx
Qwen Text Embedding v4
qwen-text-embedding-v4
8K ctx
Qwen Turbo
qwen-turbo
131K ctx
Qwen VL Max
qwen-vl-max
131K ctx
Qwen VL Plus
qwen-vl-plus
131K ctx
QwQ Plus
qwq-plus
131K ctx

V2.6
v2.6-image-to-video-flash
Wan 2.6 image-to-video flash model.

Wan
wan-v2.7-pro-edit
Edit and transform images using text instructions with the WAN 2.7 Pro model for precise, professional-grade image modifications.

Wan
wan-v2.7-edit
Transform and edit existing images with text-guided instructions using the WAN 2.7 model for creative image manipulation.

Wan
wan-v2.2-a14b-text-to-video-turbo
Wan-2.2 turbo text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text…

Wan
wan-v2.2-a14b-video-to-video
Wan-2.2 video-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts…

Wan
wan-v2.2-5b-text-to-image
Wan 2.2's 5B model generates high-resolution, photorealistic images with powerful prompt understanding and fine-grained visual detail

Wan
wan-v2.2-5b-text-to-video-fast-wan
Wan 2.2's 5B FastVideo model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding

Wan
wan-v2.7-edit-video
Wan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual…

Wan
wan-v2.2-5b-text-to-video-distill
Wan 2.2's 5B distill model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding

Wan
wan-v2.2-a14b-text-to-image
Wan 2.2's 14B model generates high-resolution, photorealistic images with powerful prompt understanding and fine-grained visual detail

Wan
wan-v2.2-a14b-image-to-video-turbo
Wan-2.2 Turbo image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text…

Wan-2.1 First-Last-Frame-to-Video
wan-flf2v
Wan-2.1 flf2v generates dynamic videos by intelligently bridging a given first frame to a desired end frame through smooth, coherent motion…

Wan-2.1 Image-to-Video
wan-i2v
Wan-2.1 is a image-to-video model that generates high-quality videos with high visual quality and motion diversity from images

Wan-2.1 Image-to-Video with LoRAs
wan-i2v-lora
Add custom LoRAs to Wan-2.1 is a image-to-video model that generates high-quality videos with high visual quality and motion diversity from…

Wan-2.1 Pro Image-to-Video
wan-pro-image-to-video
Wan-2.1 Pro is a premium image-to-video model that generates high-quality 1080p videos at 30fps with up to 6 seconds duration, delivering…

Wan-2.1 Text-to-Video
wan-t2v
Wan-2.1 is a text-to-video model that generates high-quality videos with high visual quality and motion diversity from text prompts

Wan-2.1 Text-to-Video with LoRAs
wan-t2v-lora
Add custom LoRAs to Wan-2.1 is a text-to-video model that generates high-quality videos with high visual quality and motion diversity from…

Wan 2.1 VACE Long Reframe
wan-vace-apps-long-reframe
Reframe entire videos scene-by-scene using Wan VACE 2.1

Wan-2.2 Animate Move
wan-v2.2-14b-animate-move
Wan-Animate is a video model that generates high-fidelity character videos by replicating the expressions and movements of characters from…

Wan-2.2 Animate Replace
wan-v2.2-14b-animate-replace
Wan-Animate Replace is a model that can integrate animated characters into reference videos, replacing the original character while…
Wan 2.2 Image-to-Video Flash
wan2-2-i2v-flash
Wan 2.2 Image-to-Video Plus
wan2-2-i2v-plus

Wan-2.2 Speech-to-Video 14B
wan-v2.2-14b-speech-to-video
Wan-S2V is a video model that generates high-quality videos from static images and audio, with realistic facial expressions, body…
Wan 2.2 Text-to-Image Flash
wan2-2-t2i-flash
Wan 2.2 Text-to-Image Plus
wan2-2-t2i-plus

Wan-2.2 Text-to-Video A14B with LoRAs
wan-v2.2-a14b-text-to-video-lora
Wan-2.2 text-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts.
Wan 2.2 Text-to-Video Plus
wan2-2-t2v-plus

Wan 2.2 VACE Fun A14B
wan-22-vace-fun-a14b-outpainting
VACE Fun for Wan 2.2 A14B from Alibaba-PAI

Wan 2.2 VACE Fun A14B
wan-22-vace-fun-a14b-inpainting
VACE Fun for Wan 2.2 A14B from Alibaba-PAI

Wan 2.2 VACE Fun A14B
wan-22-vace-fun-a14b-reframe
VACE Fun for Wan 2.2 A14B from Alibaba-PAI

Wan 2.2 VACE Fun A14B
wan-22-vace-fun-a14b-depth
VACE Fun for Wan 2.2 A14B from Alibaba-PAI
Wan 2.5 I2I (preview)
wan2.5-i2i-preview

Wan 2.5 Text to Image
wan-25-preview-text-to-image
Wan 2.5 text-to-image model.

Wan 2.5 Text to Video
wan-25-preview-text-to-video
Wan 2.5 text-to-video model.
Wan 2.6 Image
wan2-7-image
Wan 2.6 Image-to-Video
wan2.6-i2v
Wan 2.6 Text-to-Video
wan2.6-t2v
Wan 2.7 Image Pro
wan2-7-image-pro
Wan 2.7 Image-to-Video
wan2.7-i2v
Wan 2.7 Reference-to-Video
wan2.7-r2v
Wan 2.7 Text-to-Video
wan2.7-t2v
Wan 3
wan-3
Generate cinematic video from text with Alibaba's Wan 3.0. Supports 480p, 720p, and 1080p output.

Wan 3.0
wan-3.0
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual…

Wan 3.0
wan-3.0-reference-to-video
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual…

Wan 3.0 Prime
wan-3.0-prime
Wan 3.0 Prime Image-to-Video turns still images into dynamic, cinematic sequences with rapid turnaround, natural motion, and excellent…

Wan 3.0 Prime
alibaba/wan-3.0-prime/reference-to-video
Wan 3.0 Prime Reference-to-Video combines reference images, videos, and audio into a unified video with fast generation and strong…

Wan Effects
wan-effects
Wan Effects generates high-quality videos with popular effects from images

Wan Motion
wan-motion
Wan Motion is a streamlined character animation model that transfers motion from a driving video onto a reference character image.

Wan Text to Video
wan-v2.7-text-to-video
Wan 2.7 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual…

Wan Text to Video
alibaba/wan-3.0/text-to-video
Wan 3.0 is the latest generation AI video model, delivering enhanced motion smoothness, superior scene fidelity, and greater visual…

Wan v2.2 5B
wan-v2.2-5b-text-to-video
Wan 2.2's 5B model produces up to 5 seconds of video 720p at 24FPS with fluid motion and powerful prompt understanding

Wan v2.2 A14B Image-to-Video A14B with LoRAs
wan-v2.2-a14b-image-to-video-lora
Wan-2.2 image-to-video is a video model that generates high-quality videos with high visual quality and motion diversity from text prompts…

Wan v2.2 A14B Text-to-Image A14B with LoRAs
wan-v2.2-a14b-text-to-image-lora
Wan 2.2's 14B model with LoRA support generates high-fidelity images with enhanced prompt alignment, style adaptability.

Wan v2.6 Image to Video
v2.6-image-to-video
Wan 2.6 image-to-video model.

Wan v2.6 Reference to Video
v2.6-reference-to-video
Wan 2.6 reference-to-video model.

Wan v2.6 Text to Image
v2.6-text-to-image
Wan 2.6 text-to-image model.

Wan v2.6 Text to Video
wan/v2.6/text-to-video
Wan 2.6 text-to-video model.

Wan VACE 14B
wan-vace-14b-reframe
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Wan VACE 14B
wan-vace-14b
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Wan VACE 14B
wan-vace-14b-pose
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Wan VACE 14B
wan-vace-14b-outpainting
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Wan VACE 14B
wan-vace-14b-depth
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Wan VACE 14B
wan-vace-14b-inpainting
VACE is a video generation model that uses a source image, mask, and video to create prompted videos with controllable sources.

Wan VACE Video Edit
wan-vace-apps-video-edit
Edit videos using plain language and Wan VACE

Z Image Base
z-image-base
Z-Image is the foundation model of the Z- Image family, engineered for good quality, robust generative diversity, broad stylistic coverage…

Z Image Base Lora
z-image-base-lora
LoRA endpoint for Z-Image, the foundation model of the Z- Image family.

Z-Image Turbo
z-image-turbo

Z Image Turbo Controlnet
z-image-turbo-controlnet
Generate images from text and edge, depth or pose images using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Z Image Turbo Controlnet Lora
z-image-turbo-controlnet-lora
Generate images from text and edge, depth or pose images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Z Image Turbo Image To Image Lora
z-image-turbo-image-to-image-lora
Generate images from text and images using custom LoRA and Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Z Image Turbo Inpaint Lora
z-image-turbo-inpaint-lora
Generate images from text, an image, a mask and custom LoRA using Z-Image Turbo, Tongyi-MAI's super-fast 6B model.

Z Image Turbo Lora
z-image-turbo-lora
Text-to-Image endpoint with LoRA support for Z-Image Turbo, a super fast text-to-image model of 6B parameters developed by Tongyi-MAI.

Z-Image Turbo Seamless Tiling
z-image-turbo-tiling
Generate seamlessly tiling photorealistic images from text using Z-Image Turbo

Z-Image Turbo Seamless Tiling Lora
z-image-turbo-tiling-lora
Generate seamlessly tiling photorealistic images from text using Z-Image Turbo and custom LoRA