← All models
Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. 262,144 token context window, maximum output of 32,768 tokens. Higher uptime with 2 providers. Includes independent benchmarks from Artificial Analysis.
chatstreamingvisiontoolsjson mode
Details
- Accepts
- text
- Context window
- 131K tokens
Pricing
- Input
- 14.63 cr / 1M tokens
- Output
- 56.88 cr / 1M tokens
Prices in credits (1 credit = $0.01).
Data schema
approximateInput
| Field | Type | Description |
|---|---|---|
| promptrequired | string | The text prompt. |
Output
| Field | Type | Description |
|---|---|---|
| text | string | Generated text (choices[0].message.content). |