Poly Logo

Polylabs

Free ToolsBlog
Qwen
Qwen

Qwen3 VL 32B Instruct

Updated: September 2026

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

Specifications

Context
131K
Input Modalities
Text, Image
Output Modalities
Text
Input
$0.104/M
Output
$0.416/M

Capabilities

VISIONTEXTWEBTHINKINGWRITING

Similarly Priced Models

ModelProviderContextInput PriceOutput Price
GPT-5.6 Luna
OpenAIOpenAI
1M$0.1/M$0.6/M
GPT-5.6 Luna
OpenAIOpenAI
1M$0.1/M$0.6/M
GPT-5.6 Luna
OpenAIOpenAI
1M$0.1/M$0.6/M
GPT-5.6 Luna
OpenAIOpenAI
1M$0.1/M$0.6/M
Qwen3.5-9B
QwenQwen
262K$0.1/M$0.15/M

Performance Metrics

Intelligence Index

01

11.0

> 16% OF MODELS

Average Response Performance

Output Speed
56.3 tok/s
Time To First Token
2.68s
Time To First Answer Token
2.68s
End To End Response Time
11.56s

DATA SOURCE: Artificial Analysis

Curious about Qwen3 VL 32B Instruct?