Poly Logo

Polylabs

Free ToolsBlog
Qwen
Qwen

Qwen3 VL 32B Instruct

Updated: August 2026

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

Specifications

Context
131K
Input
$0.104/M
Output
$0.416/M

Capabilities

VISIONTEXTWEBTHINKINGWRITING

Similarly Priced Models

ModelProviderContextInput PriceOutput Price
GPT-5.6 Luna Pro
OpenAIOpenAI
1M$0.1/M$0.6/M
GPT-5.6 Luna
OpenAIOpenAI
1M$0.1/M$0.6/M
Gemma 4 31B
GoogleGoogle
262K$0.1/M$0.34/M
Qwen3.5-9B
QwenQwen
262K$0.1/M$0.15/M
ByteDance Seed: Seed-2.0-Mini
ByteDanceByteDance
262K$0.1/M$0.4/M

Performance Metrics

Intelligence Index

01

11.1

> 22% OF MODELS

Average Response Performance

Output Speed
46.5 tok/s
Time To First Token
3.32s
Time To First Answer Token
3.32s
End To End Response Time
14.08s

DATA SOURCE: Artificial Analysis

Curious about Qwen3 VL 32B Instruct?