← All models
CSM (Conversational Speech Model) is a speech generation model from Sesame that generates RVQ audio codes from text and audio inputs.
Details
- Accepts
- text
Pricing
- Price
- 0.00375 cr / character
Prices in credits (1 credit = $0.01).
Data schema
Input
| Field | Type | Description |
|---|---|---|
| scenerequired | array | The text to generate an audio from. |
| context | array | The context to generate an audio from. |
Output
| Field | Type | Description |
|---|---|---|
| audio | — | The generated audio. |