Qwen (Alibaba) logo

Qwen: Qwen3 VL 32B Instruct

qwen/qwen3-vl-32b-instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

Modalities

TextImageText

In / out price

$0.1 / $0.42 per 1M

Context

131K

Released

Oct 23, 2025

Why use Qwen3 VL 32B Instruct

  • Understands images
  • Tool and function calling (agent-ready)
  • Structured (JSON) outputs

Pricing

Input$0.1 / 1M tokens
Output$0.42 / 1M tokens

List prices via OpenRouter, checked daily. Real cost depends on the provider and on reasoning tokens.

Cost calculator

Pick a task, set how often you run it, and compare what it costs on each model.

ModelPer runPer monthPer 1,000 runs
Qwen: Qwen3 VL 32B Instruct$0.0009$0.915$0.915

Estimates from list prices via OpenRouter; real bills vary by provider and reasoning tokens.

Will it fit? Context window simulator

Each square is about one page (500 words). Colored squares are your content; grey squares are free space in the model's context window.

Your content

  • Novel (90,000 words)

Total: about 120,000 tokens (≈ 90,226 words). Token counts are estimates.

Models

Qwen: Qwen3 VL 32B Instruct

131,072 tokens ≈ 198 pages

Fits: uses 92% of the window.

Write with the best model for the job

WordGPT picks and switches models for you: drafting, editing and files in one place. Free plan, no credit card.

Try it free

More from Qwen (Alibaba)