← All models
GLM 5.3 Prime
GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. 1,000,000 token context window, maximum output of 131,072 tokens.
chatstreamingtoolsjson mode
Details
- Accepts
- text
- Context window
- 1M tokens
Pricing
- Input
- 350 cr / 1M tokens
- Output
- 1100 cr / 1M tokens
- Cached input
- 70 cr / 1M tokens
Prices in credits (1 credit = $0.01).
Data schema
approximateInput
| Field | Type | Description |
|---|---|---|
| promptrequired | string | The text prompt. |
Output
| Field | Type | Description |
|---|---|---|
| text | string | Generated text (choices[0].message.content). |