Audio models
106 models. Compare text-to-speech and speech-to-text models — voice generation and transcription run in opposite directions and are separate catalogue categories.
Price ranks a model against others of the same modality — a second of video generation and a second of image generation aren't the same unit of work, so thirds are computed within each modality, not across all of them.
106 of 106 models

ACE Step Audio Inpaint
Ace Step · Text to Speech
Modify a portion of provided audio with lyrics and/or style using ACE-Step
0.025 cr / second
Try in Studio →
ACE Step Audio Outpaint
Ace Step · Text to Speech
Extend the beginning or end of provided audio with lyrics and/or style using ACE-Step
0.025 cr / second
Try in Studio →
ACE Step Audio To Audio
Ace Step · Text to Speech
Generate music from a lyrics and example audio using ACE-Step
0.025 cr / second
Try in Studio →
Async Text to Speech Pro V1.0
Async · Text to Speech
Generate professional-quality voiceovers in seconds with Async TTS Pro model text-based control over pauses, emphasis, and timing.
0.100 cr / second
Try in Studio →
Bytedance Seed Speech Text to Speech
ByteDance · Text to Speech
Seed Speech developed by ByteDance, is a family of large-scale text-to-speech models capable of synthesizing speech that is virtually…
0.00375 cr / character
Try in Studio →
Chatterbox
Resemble AI · Text to Speech
Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.
0.00313 cr / character
Try in Studio →
Chatterbox
Resemble AI · Text to Speech
Whether you're working on memes, videos, games, or AI agents, Chatterbox brings your content to life. Use the first tts from resemble ai.
0.00313 cr / character
Try in Studio →
Chatterboxhd
Resemble AI · Text to Speech
Generate expressive, natural speech with Resemble AI's Chatterbox.
0.156 cr / second
Try in Studio →
Chatterbox Pro
Resemble AI · Text to Speech
Chatterbox (pro version), Resemble AI's first production-grade open source TTS model.
0.00500 cr / character
Try in Studio →
Cohere Transcribe
Cohere Transcribe · Speech to Text
Cohere Transcribe turns your business audio into accurate text, ready for search, analytics, and automation
0.00868 cr / second
Try in Studio →
DeepFilterNet 3
Deepfilternet3 · Text to Speech
Enhance speech audio by removing background noise and upsampling to 48KHz
0.125 cr / second
Try in Studio →
Demucs
Demucs · Text to Speech
SOTA stemming model for voice, drums, bass, guitar and more.
0.087 cr / second
Try in Studio →
Dia
Nari Labs · Text to Speech
Dia directly generates realistic dialogue from transcripts. Audio conditioning enables emotion control.
0.00500 cr / character
Try in Studio →
Dia Tts
Nari Labs · Text to Speech
Clone dialog voices from a sample audio and generate dialogs from text prompts using the Dia TTS which leverages advanced AI techniques to…
0.00500 cr / character
Try in Studio →
ElevenLabs Speech to Text
Elevenlabs · Speech to Text
Generate text from speech using ElevenLabs advanced speech-to-text model.
3.75 cr / minute
Try in Studio →
ElevenLabs Speech to Text - Scribe V2
Elevenlabs · Speech to Text
Use Scribe-V2 from ElevenLabs to do blazingly fast speech to text inferences!
1 cr / minute
Try in Studio →
ElevenLabs TTS Turbo v2.5
Elevenlabs · Text to Speech
Generate high-speed text-to-speech audio using ElevenLabs TTS Turbo v2.5.
0.00625 cr / character
Try in Studio →
ElevenLabs Voice Changer
Elevenlabs · Text to Speech
Change the voices in your audios with voices in ElevenLabs!
37.5 cr / minute
Try in Studio →
FFmpeg API [Merge Audios]
infery.ai · Text to Speech
Merge audios into a single audio using FFmpeg API!
0.021 cr / second
Try in Studio →
Flash V2.5
Elevenlabs · Text to Speech
Ultra-fast text to speech with ~75ms latency in 32 languages.
0.00625 cr / character
Try in Studio →GPT-4o Mini Transcribe
OpenAI · Speech to Text
Speech-to-text model powered by GPT-4o mini
156 cr / 1M tokens · 625 cr / 1M tokens
Try in Studio →GPT-4o Transcribe
OpenAI · Speech to Text
Speech-to-text model powered by GPT-4o
313 cr / 1M tokens · 1250 cr / 1M tokens
Try in Studio →GPT-4o Transcribe + Diarization
OpenAI · Speech to Text
313 cr / 1M tokens · 1250 cr / 1M tokens
Try in Studio →
Index TTS 2.0
Index Tts 2 · Text to Speech
Generate natural, clear speeches using Index TTS 2.0 from IndexTeam
0.250 cr / second
Try in Studio →
Inworld TTS-1.5 Max
Inworld Tts · Text to Speech
Text to Speech Endpoint for Inworld's TTS-1.5 Max.
0.00125 cr / character
Try in Studio →
Kling TTS
Kling · Text to Speech
Generate speech from text prompts and different voices using the Kling TTS model, which leverages advanced AI techniques to create…
0.875 cr / request
Try in Studio →
Kling Video
Kling · Text to Speech
Generate audio from input videos using Kling
4.38 cr / clip
Try in Studio →
Kling Video Create Voice
Kling · Text to Speech
Create Voices to be used with Kling Models Voice Control
0.875 cr / request
Try in Studio →
Maya
Maya · Text to Speech
Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise…
0.250 cr / second
Try in Studio →
Maya
Maya · Text to Speech
Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise…
0.250 cr / second
Try in Studio →
Maya1
Maya · Text to Speech
Maya1 is a state-of-the-art speech model by Maya Research for expressive voice generation, built to capture real human emotion and precise…
0.250 cr / second
Try in Studio →
Minimax
MiniMax · Text to Speech
Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to…
0.013 cr / character
Try in Studio →
Minimax
MiniMax · Text to Speech
Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques…
0.00750 cr / character
Try in Studio →
MiniMax Speech-02 HD
MiniMax · Text to Speech
Generate speech from text prompts and different voices using the MiniMax Speech-02 HD model, which leverages advanced AI techniques to…
0.013 cr / character
Try in Studio →
MiniMax Speech-02 Turbo
MiniMax · Text to Speech
Generate fast speech from text prompts and different voices using the MiniMax Speech-02 Turbo model, which leverages advanced AI techniques…
0.00750 cr / character
Try in Studio →
MiniMax Speech 2.6 [HD]
MiniMax · Text to Speech
Generate speech from text prompts and different voices using the MiniMax Speech-2.6 HD model, which leverages advanced AI techniques to…
0.013 cr / character
Try in Studio →
MiniMax Speech 2.6 [Turbo]
MiniMax · Text to Speech
Generate speech from text prompts and different voices using the MiniMax Speech-2.6 HD model, which leverages advanced AI techniques to…
0.00750 cr / character
Try in Studio →
MiniMax Speech 2.8 [HD]
MiniMax · Text to Speech
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 HD model, which leverages advanced AI techniques to…
0.013 cr / character
Try in Studio →
MiniMax Speech 2.8 [Turbo]
MiniMax · Text to Speech
Generate speech from text prompts and different voices using the MiniMax Speech-2.8 Turbo model, which leverages advanced AI techniques to…
0.00750 cr / character
Try in Studio →
Mirelo SFX
Mirelo AI · Text to Speech
Generate synced sounds for any video, and return the new sound track (like MMAudio)
0.156 cr / second
Try in Studio →
Mirelo SFX1.6
Mirelo AI · Text to Speech
Erase and replace any moment in your audio with AI-driven precision.
0.156 cr / second
Try in Studio →
Mirelo SFX1.6
Mirelo AI · Text to Speech
Extend any sound effect with seamless, natural tails.
0.156 cr / second
Try in Studio →
Mirelo SFX V1.5
Mirelo AI · Text to Speech
Generate synced sounds for any video, and return the new sound track (like MMAudio)
0.156 cr / second
Try in Studio →
Nemotron 3 Nano Omni
NVIDIA · Speech to Text
Audio reasoning variant of NVIDIA's Nemotron 3 Nano Omni.
1250 cr / 1M tokens
Try in Studio →
Nemotron Asr Multilingual
NVIDIA · Speech to Text
Nemotron-ASR-Streaming is a multi lingual, streaming Automatic Speech Recognition (ASR) engineered to deliver high-quality multi lingual…
0.139 cr / second
Try in Studio →
Orpheus TTS
Canopy Labs · Text to Speech
Orpheus TTS is a state-of-the-art, Llama-based Speech-LLM designed for high-quality, empathetic text-to-speech generation.
0.00625 cr / character
Try in Studio →Audio, on Infery.ai
Audio splits into two catalogue categories. audio_stt is speech-to-text: it transcribes an audio file into text. audio_tts is named for text-to-speech, and about half of its models do exactly that — turn a written script into a voice — but the other half take an existing audio or video clip as input instead, generating sound effects, foley timed to a video, or extending and filling an existing clip. That audio-to-audio work sits inside the same catalogue category as voice generation, not a separate one.
It is a catalogue, not one model
Within audio_tts, primary input is the real split: a text-input model turns a script into a voice, usually billed per character; an audio- or video-input model instead transforms existing sound, usually billed per image or per second. audio_stt models differ in maximum input length and language coverage, listed on each model's schema. Narrating a script and transcribing a recording are opposite jobs done by different models — audio_tts alone does not do both.
Open the Studio with a audio model already picked
Free trial credits, no card. Filter above, or start from a blank chat and pick as you go.


