Qwen: Qwen3 VL 235B A22B Thinking✓ Catalog verified
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math. The series emphasizes robust perception (recognition of diverse real-world and synthetic categories), spatial understanding (2D/3D grounding), and long-form visual comprehension, with competitive results on public multimodal benchmarks for both perception and reasoning. Beyond analysis, Qwen3-VL supports agentic interaction and tool use: it can follow complex instructions over multi-image, multi-turn dialogues; align text to video timelines for precise temporal queries; and operate GUI elements for automation tasks. The models also enable visual coding workflows, turning sketches or mockups into code and assisting with UI debugging, while maintaining strong text-only performance comparable to the flagship Qwen3 language models. This makes Qwen3-VL suitable for production scenarios spanning document AI, multilingual OCR, software/UI assistance, spatial/embodied tasks, and research on vision-language agents.
Overview
Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math. The series emphasizes robust perception (recognition of diverse real-world and synthetic categories), spatial understanding (2D/3D grounding), and long-form visual comprehension, with competitive results on public multimodal benchmarks for both perception and reasoning. Beyond analysis, Qwen3-VL supports agentic interaction and tool use: it can follow complex instructions over multi-image, multi-turn dialogues; align text to video timelines for precise temporal queries; and operate GUI elements for automation tasks. The models also enable visual coding workflows, turning sketches or mockups into code and assisting with UI debugging, while maintaining strong text-only performance comparable to the flagship Qwen3 language models. This makes Qwen3-VL suitable for production scenarios spanning document AI, multilingual OCR, software/UI assistance, spatial/embodied tasks, and research on vision-language agents.
Access: available through the official Qwen API. Context window: 131K.
Specifications
| API identifier | qwen/qwen3-vl-235b-a22b-thinking |
| Provider | Qwen |
| Model type | Multimodal LLM |
| Context window | 131K catalog |
| Input modalities | text · image |
| Output modalities | text |
| Released | Sep 23, 2025 |
| Tokenizer | Qwen3 |
| Moderated | No |
| Architecture modality | text+image->text |
API defaults
{
"top_k": 20,
"top_p": 0.95,
"temperature": 0.8,
"presence_penalty": null,
"frequency_penalty": null,
"repetition_penalty": 1
}Pricing
Live pricing components for qwen/qwen3-vl-235b-a22b-thinking as published in the catalog.
| Pricing component | Raw unit price | Normalized | Unit |
|---|---|---|---|
| Prompt tokens | $2.6e-7 | $0.26 / 1M | Per input token |
| Completion tokens | $0.0000026 | $2.60 / 1M | Per output token |
| Request fee | N/A | N/A | Per request |
| Image fee | N/A | N/A | Per image unit |
| Web search fee | N/A | N/A | Per search request |
Cost calculator
Estimate monthly spend from your own token volumes.
Estimated Total Cost
$0.65Calculated from current list pricing in the ModelsAtlas catalog.Capabilities
| Capability | Status | What it means |
|---|---|---|
| Visual Understanding | ✓ Supported | Image and document analysis support |
| Audio Processing | Not advertised | Speech and voice aligned flows |
| Tool Calling | ✓ Supported | Supports tools / function calling |
| Self-Hosting | ✓ Supported | Deploy outside managed APIs |
Capability signals
Directional signals derived from model metadata and capability tags — not official benchmark submissions.
Quick start
Call Qwen: Qwen3 VL 235B A22B Thinking through an OpenAI-compatible client.
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key="YOUR_API_KEY"
)
response = client.chat.completions.create(
model="qwen/qwen3-vl-235b-a22b-thinking",
messages=[{"role": "user", "content": "Explain quantum physics."}]
)
print(response.choices[0].message.content)Alternatives
Closest models by context window from a different provider.
| Model | Provider | Context | Input /1M | Actions |
|---|---|---|---|---|
| AionLabs: Aion-1.0 | Aion Labs | 131K | $4.00 | View → |
| AionLabs: Aion-1.0-Mini | Aion Labs | 131K | $0.70 | View → |
Sources & attribution
Pricing and metadata are maintained in the ModelsAtlas catalog.
Frequently asked questions
How much does Qwen: Qwen3 VL 235B A22B Thinking cost?
$0.26 per 1M input tokens and $2.60 per 1M output tokens on the official Qwen API.
What is the context window of Qwen: Qwen3 VL 235B A22B Thinking?
Qwen: Qwen3 VL 235B A22B Thinking supports up to 131K tokens of context.
Does Qwen: Qwen3 VL 235B A22B Thinking support tool / function calling?
Yes, Qwen: Qwen3 VL 235B A22B Thinking supports tool / function calling — you can register tools and the model will emit structured tool calls.
How do I access Qwen: Qwen3 VL 235B A22B Thinking?
Use the official Qwen API with the model id `qwen/qwen3-vl-235b-a22b-thinking`. See the Quick start section above for code examples.