← All models
Stable Audio 3 Medium Audio Outpainting
stable-audio-3-medium-audio-outpainting
Text to Speechby Stability AI
Stable Audio 3 Medium audio outpainting is a 1.4 billion parameter latent diffusion model that extends existing stereo audio beyond its original endpoint via causal continuation guided by text prompts.
Details
- Accepts
- audio
Pricing
- Price
- 5.58 cr / clip
Prices in credits (1 credit = $0.01).
Data schema
Input
| Field | Type | Description |
|---|---|---|
| seed | — | Random seed for reproducible outputs. Omit for a random seed. |
| promptrequired | string | Text description guiding the extension content. |
| bitrate | string | Audio bitrate for compressed output formats (e.g., mp3, aac, opus). Format e.g. '192k' or '320k'. Ignored for lossless formats (wav, flac). |
| audio_urlrequired | string | Source audio to extend. |
| sync_mode | boolean | If True, the audio is returned inline as a data URI and the result is not saved to the request history. |
| output_format | stringenum: mp3, wav, flac, ogg, opus, m4a… | Container format for the generated audio output. |
| guidance_scale | number | Classifier-free guidance scale. Higher values follow the prompt more strictly. Only effective on base (non-distilled) checkpoints. |
| negative_prompt | string | Text description of qualities to avoid in the output. |
| num_inference_steps | integer | Number of sampling steps. Post-trained (distilled) checkpoints look good with the default 8 and gain little from going higher. |
| extend_seconds_after | number | Seconds of new audio to generate after the source audio. |
| enable_safety_checker | boolean | Enable NSFW content safety checking. |
| extend_seconds_before | number | Seconds of new audio to generate before the source audio. |
| enable_prompt_expansion | boolean | If True, the prompt will be expanded using an LLM for more detailed and higher quality results. |
Output
| Field | Type | Description |
|---|---|---|
| seed | integer | The random seed used for generation. |
| audio | — | The generated audio clip. |
| prompt | string | The prompt used for generation. |