infery
← All models

CSM-1B

csm-1b

Musicby Sesame

CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.

Details

Accepts
text

Pricing

Price
0.00375 cr / character

Prices in credits (1 credit = $0.01).

Data schema

Input

FieldTypeDescription
scenerequiredarrayThe text to generate an audio from.
contextarrayThe context to generate an audio from.

Output

FieldTypeDescription
audioThe generated audio.