infery
← All modalities

Audio models

106 models. Compare text-to-speech and speech-to-text models — voice generation and transcription run in opposite directions and are separate catalogue categories.

Price ranks a model against others of the same modality — a second of video generation and a second of image generation aren't the same unit of work, so thirds are computed within each modality, not across all of them.

Price
Capability

106 of 106 models

What it is

Audio, on Infery.ai

Audio splits into two catalogue categories. audio_stt is speech-to-text: it transcribes an audio file into text. audio_tts is named for text-to-speech, and about half of its models do exactly that — turn a written script into a voice — but the other half take an existing audio or video clip as input instead, generating sound effects, foley timed to a video, or extending and filling an existing clip. That audio-to-audio work sits inside the same catalogue category as voice generation, not a separate one.

How the models differ

It is a catalogue, not one model

Within audio_tts, primary input is the real split: a text-input model turns a script into a voice, usually billed per character; an audio- or video-input model instead transforms existing sound, usually billed per image or per second. audio_stt models differ in maximum input length and language coverage, listed on each model's schema. Narrating a script and transcribing a recording are opposite jobs done by different models — audio_tts alone does not do both.

Open the Studio with a audio model already picked

Free trial credits, no card. Filter above, or start from a blank chat and pick as you go.