infery
← All models

Qwen3 VL 30B A3B Instruct

qwen3-vl-30b-a3b-instruct

Chat & Textby Qwen

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. 262,144 token context window, maximum output of 32,768 tokens. Higher uptime with 5 providers. Includes independent benchmarks from Artificial Analysis.

chatstreamingvisiontoolsjson mode

Details

Accepts
text
Context window
131K tokens

Pricing

Input
18.75 cr / 1M tokens
Output
75 cr / 1M tokens

Prices in credits (1 credit = $0.01).

Data schema

approximate

Input

FieldTypeDescription
promptrequiredstringThe text prompt.

Output

FieldTypeDescription
textstringGenerated text (choices[0].message.content).