infery
← All models

Z.AI models

16 models on Infery.ai.

GLM 4.5

GLM 4.5

glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications.

131K ctx

75 cr in / 1M275 cr out / 1M13.75 cr cached in / 1M
ChatStreamingToolsJSON
GLM 4.5 Air

GLM 4.5 Air

glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications.

131K ctx

16.25 cr in / 1M106 cr out / 1M3.13 cr cached in / 1M
ChatStreamingToolsJSON
GLM 4.5V

GLM 4.5V

glm-4.5v

GLM-4.5V is a vision-language foundation model for multimodal agent applications.

66K ctx

75 cr in / 1M225 cr out / 1M13.75 cr cached in / 1M
ChatStreamingVisionToolsJSON
GLM 4.6

GLM 4.6

glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from…

203K ctx

53.75 cr in / 1M219 cr out / 1M10 cr cached in / 1M
ChatStreamingToolsJSON
GLM 4.6V

GLM 4.6V

glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents…

131K ctx

37.5 cr in / 1M113 cr out / 1M6.88 cr cached in / 1M
ChatStreamingVisionToolsJSON
GLM 4.7

GLM 4.7

glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step…

203K ctx

50 cr in / 1M219 cr out / 1M10 cr cached in / 1M
ChatStreamingToolsJSON
GLM 4.7 Flash

GLM 4.7 Flash

glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency.

203K ctx

7.56 cr in / 1M50 cr out / 1M
ChatStreamingToolsJSON
GLM 5

GLM 5

glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.

203K ctx

75 cr in / 1M240 cr out / 1M15 cr cached in / 1M
ChatStreamingToolsJSON
GLM 5.1

GLM 5.1

glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks.

203K ctx

121 cr in / 1M380 cr out / 1M22.43 cr cached in / 1M
ChatStreamingToolsJSON
GLM 5.2

GLM 5.2

glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. 1,048,576 token context window, maximum output of 32,768 tokens.

1M ctx

175 cr in / 1M550 cr out / 1M17.5 cr cached in / 1M
ChatStreamingToolsJSON
GLM 5.3

GLM 5.3

glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks.

1M ctx

175 cr in / 1M550 cr out / 1M32.5 cr cached in / 1M
ChatStreamingToolsJSON
GLM 5.3 Flash

GLM 5.3 Flash

glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. 1,048,576 token context window, maximum output of 131,072 tokens.

1M ctx

11.25 cr in / 1M37.5 cr out / 1M2.25 cr cached in / 1M
ChatStreamingVisionToolsJSON
GLM 5 Turbo

GLM 5 Turbo

glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw…

203K ctx

150 cr in / 1M500 cr out / 1M30 cr cached in / 1M
ChatStreamingToolsJSON
GLM 5V Turbo

GLM 5V Turbo

glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks.

203K ctx

150 cr in / 1M500 cr out / 1M30 cr cached in / 1M
ChatStreamingVisionToolsJSON
GLM Flash Latest

GLM Flash Latest

glm-flash-latest

This model always redirects to the latest model in the GLM Flash family. 1,310,720 token context window, maximum output of 131,072 tokens.

1M ctx

9.38 cr in / 1M31.25 cr out / 1M1.88 cr cached in / 1M
ChatStreamingVisionToolsJSON
GLM Latest

GLM Latest

glm-latest

This model always redirects to the latest GLM model from Z.ai. 1,048,576 token context window, maximum output of 131,072 tokens.

1M ctx

113 cr in / 1M375 cr out / 1M18.75 cr cached in / 1M
ChatStreamingToolsJSON