← All models
Stable Audio 3 Medium is a 1.4 billion parameter latent diffusion model that generates high-quality stereo music up to 6 minutes from text prompts, trained on fully licensed data for safe commercial use.
Details
- Accepts
- text
Pricing
- Price
- 4.7 cr / track
Prices in credits (1 credit = $0.01).
Data schema
Input
| Field | Type | Description |
|---|---|---|
| seed | — | Random seed for reproducible outputs. Omit for a random seed. |
| promptrequired | string | Text description of the audio to generate. |
| bitrate | string | Audio bitrate for compressed output formats (e.g., mp3, aac, opus). Format e.g. '192k' or '320k'. Ignored for lossless formats (wav, flac). |
| duration | number | Duration of the generated audio in seconds. The medium model supports up to 380 seconds (~6m20s). |
| sync_mode | boolean | If True, the audio is returned inline as a data URI and the result is not saved to the request history. |
| output_format | stringenum: mp3, wav, flac, ogg, opus, m4a… | Container format for the generated audio output. |
| guidance_scale | number | Classifier-free guidance scale. Higher values follow the prompt more strictly. Only effective on base (non-distilled) checkpoints. |
| negative_prompt | string | Text description of qualities to avoid in the output. |
| num_inference_steps | integer | Number of sampling steps. Post-trained (distilled) checkpoints look good with the default 8 and gain little from going higher. |
| enable_safety_checker | boolean | Enable NSFW content safety checking. |
| enable_prompt_expansion | boolean | If True, the prompt will be expanded using an LLM for more detailed and higher quality results. |
Output
| Field | Type | Description |
|---|---|---|
| seed | integer | The random seed used for generation. |
| audio | — | The generated audio clip. |
| prompt | string | The prompt used for generation. |