Qwen3 VL 32B Instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video.
text + image → text · made by Alibaba
$0.1→$0.42per Mtok in / out
Nous Research · 5 sellers
Sold by 5 ways
| Seller | Lane | Rate |
|---|---|---|
| Nous Research aggregator | standard | $0.1per Mtok in$0.42per Mtok outinference-api.nousresearch.com · read 2026-09-12 |
| OpenRouter aggregator | standard | $0.1per Mtok in$0.42per Mtok outopenrouter.ai · read 2026-08-24 |
| SiliconFlow aggregator | standard | $0.2per Mtok in$0.6per Mtok outmodels.dev · read 2026-09-16 |
| Together AI aggregator | standard | $0.5per Mtok in$1.5per Mtok outraw.githubusercontent.com · read 2026-09-16 |
| Fireworks AI aggregator | standard | $0.9per Mtok in$0.9per Mtok outraw.githubusercontent.com · read 2026-09-16 |
About
Qwen3 VL 32B Instruct — a text model from Alibaba, sold by 5 companies from $0.1 in and $0.42 out per million tokens.
It takes text and images and returns text, with a context window of 131,072 tokens. It was published in October 2025. Its sellers say it can call a tool. The catalogue files it under chat. Five companies sell it. The cheapest is $0.1 in and $0.42 out per million tokens at Nous Research.
Every current figure
- Maker
- Alibaba
- Register
- model
- Takes
- text + image
- Returns
- text
- Context
- 131,072 tokens
- Published
- October 2025
- Parameters
- 33 billion
- Licence
- not read
- Sellers
- 5
- Maker's own price
- not read
- Price
- $0.1 in and $0.42 out per million tokens — Nous Research
Known as 1 name
qwen/qwen3-vl-32b-instruct