infery
← All models

Stability AI models

34 models on Infery.ai.

High Quality Stable Video Diffusion

stable-video

Generate short video clips from your images using SVD v1.1

9.38 cr / video
Video output

Stable Audio 2.5

stable-audio-25-text-to-audio

Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI

25 cr / track
Music output

Stable Audio 2.5

stable-audio-25-audio-to-audio

Generate high quality music and sound effects using Stable Audio 2.5 from StabilityAI

25 cr / clip
Audio output

Stable Audio 3

stable-audio-3-small-music-audio-to-audio

Stable Audio 3 Small Music audio-to-audio is a 459 million parameter latent diffusion model that transforms input music into new variations…

3.1 cr / clip
Audio output

Stable Audio 3

stable-audio-3-medium-text-to-audio

Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text…

4.7 cr / track
Music output

Stable Audio 3

stable-audio-3-small-music-base-text-to-audio

Stable Audio 3 Small Music Base is the foundational 459 million parameter checkpoint generating full music compositions up to 2 minutes…

3.51 cr / track
Music output

Stable Audio 3

stable-audio-3-small-sfx-audio-to-audio

Stable Audio 3 Small SFX audio-to-audio is a 459 million parameter latent diffusion model that transforms input audio into new sound-effect…

3 cr / clip
Audio output

Stable Audio 3 Medium Audio Inpainting

stable-audio-3-medium-audio-inpainting

Stable Audio 3 Medium audio inpainting is a 1.4 billion parameter latent diffusion model that fills in or reworks selected segments of a…

5.53 cr / clip
Audio output

Stable Audio 3 Medium Audio Outpainting

stable-audio-3-medium-audio-outpainting

Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its…

5.58 cr / clip
Audio output

Stable Audio 3 Medium Audio to Audio

stable-audio-3-medium-audio-to-audio

Stable Audio 3 Medium audio-to-audio is a 1.4 billion parameter latent diffusion model that transforms an input audio clip into new stereo…

5.21 cr / clip
Audio output

Stable Audio 3 Medium Base Audio Inpainting

stable-audio-3-medium-base-audio-inpainting

Stable Audio 3 Medium Base audio inpainting is the foundational 1.4 billion parameter checkpoint for editing or filling selected stereo…

6.71 cr / clip
Audio output

Stable Audio 3 Medium Base Audio Outpainting

stable-audio-3-medium-base-audio-outpainting

Stable Audio 3 Medium Base audio outpainting is the foundational 1.4 billion parameter checkpoint that extends existing stereo audio with…

6.92 cr / clip
Audio output

Stable Audio 3 Medium Base Audio to Audio

stable-audio-3-medium-base-audio-to-audio

Stable Audio 3 Medium Base audio-to-audio is the foundational 1.4 billion parameter checkpoint that transforms input audio into new stereo…

6.54 cr / clip
Audio output

Stable Audio 3 Medium Base Text to Audio

stable-audio-3-medium-base-text-to-audio

Stable Audio 3 Medium Base is the foundational 1.4 billion parameter text-to-audio checkpoint generating stereo music up to 6 minutes…

5.99 cr / track
Music output

Stable Audio 3 Small Music Audio Inpainting

stable-audio-3-small-music-audio-inpainting

Stable Audio 3 Small Music audio inpainting is a 459 million parameter latent diffusion model that fills in or reworks selected segments of…

3.25 cr / clip
Audio output

Stable Audio 3 Small Music Audio Outpainting

stable-audio-3-small-music-audio-outpainting

Stable Audio 3 Small Music audio outpainting is a 459 million parameter latent diffusion model that extends music compositions beyond their…

3.31 cr / clip
Audio output

Stable Audio 3 Small Music Base Audio Inpainting

stable-audio-3-small-music-base-audio-inpainting

Stable Audio 3 Small Music Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected music…

4.16 cr / clip
Audio output

Stable Audio 3 Small Music Base Audio Outpainting

stable-audio-3-small-music-base-audio-outpainting

Stable Audio 3 Small Music Base audio outpainting is the foundational 459 million parameter checkpoint that extends music tracks via causal…

4.24 cr / clip
Audio output

Stable Audio 3 Small Music Base Audio to Audio

stable-audio-3-small-music-base-audio-to-audio

Stable Audio 3 Small Music Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input music into new…

4 cr / clip
Audio output

Stable Audio 3 Small Music Text to Audio

stable-audio-3-small-music-text-to-audio

Stable Audio 3 Small Music is a 459 million parameter latent diffusion model that generates full stereo music compositions up to 2 minutes…

2.71 cr / track
Music output

Stable Audio 3 Small SFX Audio Inpainting

stable-audio-3-small-sfx-audio-inpainting

Stable Audio 3 Small SFX audio inpainting is a 459 million parameter latent diffusion model that fills in or reworks selected segments of a…

3.24 cr / clip
Audio output

Stable Audio 3 Small SFX Audio Outpainting

stable-audio-3-small-sfx-audio-outpainting

Stable Audio 3 Small SFX audio outpainting is a 459 million parameter latent diffusion model that extends sound-effect tracks beyond their…

3.39 cr / clip
Audio output

Stable Audio 3 Small SFX Base Audio Inpainting

stable-audio-3-small-sfx-base-audio-inpainting

Stable Audio 3 Small SFX Base audio inpainting is the foundational 459 million parameter checkpoint for editing or filling selected…

4.14 cr / clip
Audio output

Stable Audio 3 Small SFX Base Audio Outpainting

stable-audio-3-small-sfx-base-audio-outpainting

Stable Audio 3 Small SFX Base audio outpainting is the foundational 459 million parameter checkpoint that extends sound-effect tracks via…

4.21 cr / clip
Audio output

Stable Audio 3 Small SFX Base Audio to Audio

stable-audio-3-small-sfx-base-audio-to-audio

Stable Audio 3 Small SFX Base audio-to-audio is the foundational 459 million parameter checkpoint that transforms input audio into new…

3.94 cr / clip
Audio output

Stable Audio 3 Small SFX Base Text to Audio

stable-audio-3-small-sfx-base-text-to-audio

Stable Audio 3 Small SFX Base is the foundational 459 million parameter checkpoint generating sound effects from text prompts, intended as…

3.54 cr / track
Music output

Stable Audio 3 Small SFX Text to Audio

stable-audio-3-small-sfx-text-to-audio

Stable Audio 3 Small SFX is a 459 million parameter latent diffusion model that generates high-quality sound effects from text prompts…

2.58 cr / track
Music output

Stable Audio Open

stable-audio

Open source text-to-audio model.

0.156 cr / sec
Music output

Stable Cascade

stable-cascade

Stable Cascade: Image generation on a smaller & cheaper latent space.

0.072 cr / sec
Image output

Stable Diffusion 3.5 Large

stable-diffusion-3.5-large

Stable Diffusion 3.5 Large generates high-resolution images with fine details, strong typography, and diverse artistic styles using the…

8.13 cr / image
Image output

Stable Diffusion 3.5 Large

stable-diffusion-v35-large

Stable Diffusion 3.5 Large is a Multimodal Diffusion Transformer (MMDiT) text-to-image model that features improved performance in image…

8.13 cr / MP
Image output

Stable Diffusion 3.5 Large Turbo

stable-diffusion-3.5-large-turbo

Stable Diffusion 3.5 Large Turbo generates high-resolution images in fewer steps with fine details and diverse artistic styles.

5 cr / image
Image output

Stable Diffusion 3.5 Medium

stable-diffusion-3.5-medium

Stable Diffusion 3.5 Medium is a 2.5 billion parameter image model with improved MMDiT-X architecture for high-quality, multi-resolution…

4.38 cr / image
Image output

Stable Diffusion V3

stable-diffusion-v3-medium

Stable Diffusion 3 Medium (Text to Image) is a Multimodal Diffusion Transformer (MMDiT) model that improves image quality, typography…

4.38 cr / image
Image output