infery
← All models
GLM 5.3 FlashX

GLM 5.3 FlashX

glm-5.3-flashx

Chat & Textby Z.AI

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s. 1,048,576 token context window, maximum output of 131,072 tokens.

chatstreamingvisiontoolsjson mode

Details

Accepts
text
Context window
1M tokens

Pricing

Input
46.25 cr / 1M tokens
Output
156 cr / 1M tokens
Cached input
9.38 cr / 1M tokens

Prices in credits (1 credit = $0.01).

Data schema

approximate

Input

FieldTypeDescription
promptrequiredstringThe text prompt.

Output

FieldTypeDescription
textstringGenerated text (choices[0].message.content).