infery
← All modalities

Music models

41 models. Compare AI music generation models by input — a text prompt or lyrics — and by output length and pricing unit.

Price ranks a model against others of the same modality — a second of video generation and a second of image generation aren't the same unit of work, so thirds are computed within each modality, not across all of them.

Price
Capability

41 of 41 models

What it is

Music, on Infery.ai

Music is its own catalogue category, separate from text-to-speech and speech-to-text: most of it generates an instrumental or vocal track from a text prompt. A small number also require a reference audio clip — voice cloning is the clear case, where the model needs a sample of the voice it's meant to sing in, not just a description of one.

How the models differ

It is a catalogue, not one model

Models differ in whether they accept lyrics as a direct input alongside a style description — ace-step, for one, takes both a lyrics field and a tags field, where most models take a single prompt — and in maximum track length. Pricing unit varies more here than in any other modality: per-second and per-image are tied as the two most common, but per-operation, per-character, per-request, per-minute, per-megapixel and per-token pricing all appear too — eight units in total, so no single price-per-track figure applies across the category.

Open the Studio with a music model already picked

Free trial credits, no card. Filter above, or start from a blank chat and pick as you go.