infery
← All models

Qwen3 VL 32B Instruct

qwen3-vl-32b-instruct

Chat & Textby Qwen

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. 131,072 token context window, maximum output of 32,768 tokens. Includes independent benchmarks from Artificial Analysis.

chatstreamingvisiontoolsjson mode

Details

Accepts
text
Context window
131K tokens

Pricing

Input
13 cr / 1M tokens
Output
52 cr / 1M tokens

Prices in credits (1 credit = $0.01).

Data schema

approximate

Input

FieldTypeDescription
promptrequiredstringThe text prompt.

Output

FieldTypeDescription
textstringGenerated text (choices[0].message.content).