infery
← All models
GLM 5.3 Prime

GLM 5.3 Prime

glm-5.3-prime

Chat & Textby Z.AI

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration. 1,000,000 token context window, maximum output of 131,072 tokens.

chatstreamingtoolsjson mode

Details

Accepts
text
Context window
1M tokens

Pricing

Input
350 cr / 1M tokens
Output
1100 cr / 1M tokens
Cached input
70 cr / 1M tokens

Prices in credits (1 credit = $0.01).

Data schema

approximate

Input

FieldTypeDescription
promptrequiredstringThe text prompt.

Output

FieldTypeDescription
textstringGenerated text (choices[0].message.content).