infery
← All models

ACE Step Audio Inpaint

ace-step-audio-inpaint

Text to Speechby Ace Step

Modify a portion of provided audio with lyrics and/or style using ACE-Step

Details

Accepts
audio

Pricing

Price
0.025 cr / second

Prices in credits (1 credit = $0.01).

Data schema

Input

FieldTypeDescription
seedRandom seed for reproducibility. If not provided, a random seed will be used.
tagsrequiredstringComma-separated list of genre tags to control the style of the generated audio. Can also be supplied as `prompt`.
lyricsstringLyrics to be sung in the audio. If not provided or if [inst] or [instrumental] is the content of this field, no lyrics will be sung. Use control structures like [verse], [chorus] and [bridge] to control the structure of the song.
end_timenumberend time in seconds for the inpainting process.
variancenumberVariance for the inpainting process. Higher values can lead to more diverse results.
audio_urlrequiredstringURL of the audio file to be inpainted.
schedulerstringenum: euler, heunScheduler to use for the generation process.
start_timenumberstart time in seconds for the inpainting process.
guidance_typestringenum: cfg, apg, cfg_starType of CFG to use for the generation process.
guidance_scalenumberGuidance scale for the generation.
number_of_stepsintegerNumber of steps to generate the audio.
granularity_scaleintegerGranularity scale for the generation process. Higher values can reduce artifacts.
guidance_intervalnumberGuidance interval for the generation. 0.5 means only apply guidance in the middle steps (0.25 * infer_steps to 0.75 * infer_steps)
tag_guidance_scalenumberTag guidance scale for the generation.
end_time_relative_tostringenum: start, endWhether the end time is relative to the start or end of the audio.
lyric_guidance_scalenumberLyric guidance scale for the generation.
minimum_guidance_scalenumberMinimum guidance scale for the generation after the decay.
start_time_relative_tostringenum: start, endWhether the start time is relative to the start or end of the audio.
guidance_interval_decaynumberGuidance interval decay for the generation. Guidance scale will decay from guidance_scale to min_guidance_scale in the interval. 0.0 means no decay.

Output

FieldTypeDescription
seedintegerThe random seed used for the generation process.
tagsstringThe genre tags used in the generation process.
audioThe generated audio file.
lyricsstringThe lyrics used in the generation process.