infery
← All models

Qwen3 VL 8B Instruct

qwen3-vl-8b-instruct

Chat & Textby Qwen

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. 262,144 token context window, maximum output of 32,768 tokens. Higher uptime with 2 providers. Includes independent benchmarks from Artificial Analysis.

chatstreamingvisiontoolsjson mode

Details

Accepts
text
Context window
131K tokens

Pricing

Input
14.63 cr / 1M tokens
Output
56.88 cr / 1M tokens

Prices in credits (1 credit = $0.01).

Data schema

approximate

Input

FieldTypeDescription
promptrequiredstringThe text prompt.

Output

FieldTypeDescription
textstringGenerated text (choices[0].message.content).