infery
← All models

Eleven v4

tts-eleven-v4

Text to Speechby Elevenlabs

Generate expressive speech with Eleven v4 from ElevenLabs. Control delivery with audio tags, voice stability, similarity settings, and IPA pronunciation.

Details

Accepts
text

Pricing

Price
0.010 cr / character

Prices in credits (1 credit = $0.01).

Data schema

Input

FieldTypeDescription
seed—Seed for best-effort reproducibility. Identical output is not guaranteed.
textrequiredstringThe text to convert to speech. Supports audio tags such as [whispering] and IPA pronunciation enclosed in forward slashes.
voicestringThe voice to use for speech generation
stabilitynumberVoice stability. Lower values allow more expressive delivery; higher values make delivery more consistent.
timestampsbooleanWhether to return character-level timing information with the generated audio.
language_code—Language code (ISO 639-1) for speech generation and text normalization.
output_formatstringenum: mp3_22050_32, mp3_44100_32, mp3_44100_64, mp3_44100_96, mp3_44100_128, mp3_44100_192…Output format of the generated audio. Formatted as codec_sample_rate_bitrate.
similarity_boostnumberHow closely the output follows the reference voice. Higher values increase similarity but may reduce naturalness.
apply_text_normalizationstringenum: auto, on, offWhether to normalize text such as numbers and dates before generation.

Output

FieldTypeDescription
audio—The generated audio file
timestamps—Timestamps for each word in the generated speech. Only returned if `timestamps` is set to True in the request.