Qwen: Qwen2.5-VL 7B Instruct✓ Catalog verified
Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements: - SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. - Understanding videos of 20min+: Qwen2.5-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. - Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2.5-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions. - Multilingual Support: to serve global users, besides English and Chinese, Qwen2.5-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc. For more details, see this blog post and GitHub repo. Usage of this model is subject to Tongyi Qianwen LICENSE AGREEMENT.
Overview
Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements: - SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. - Understanding videos of 20min+: Qwen2.5-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. - Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2.5-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions. - Multilingual Support: to serve global users, besides English and Chinese, Qwen2.5-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc. For more details, see this blog post and GitHub repo. Usage of this model is subject to Tongyi Qianwen LICENSE AGREEMENT.
Access: available through the official Qwen API. Context window: 33K.
Specifications
| API identifier | qwen/qwen-2.5-vl-7b-instruct |
| Provider | Qwen |
| Model type | Multimodal LLM |
| Context window | 33K catalog |
| Input modalities | text · image |
| Output modalities | text |
| Released | Aug 28, 2024 |
| Tokenizer | Qwen |
| Moderated | No |
| Architecture modality | text+image->text |
Pricing
Live pricing components for qwen/qwen-2.5-vl-7b-instruct as published in the catalog.
| Pricing component | Raw unit price | Normalized | Unit |
|---|---|---|---|
| Prompt tokens | $2e-7 | $0.20 / 1M | Per input token |
| Completion tokens | $2e-7 | $0.20 / 1M | Per output token |
| Request fee | N/A | N/A | Per request |
| Image fee | N/A | N/A | Per image unit |
| Web search fee | N/A | N/A | Per search request |
Cost calculator
Estimate monthly spend from your own token volumes.
Estimated Total Cost
$0.14Calculated from current list pricing in the ModelsAtlas catalog.Capabilities
| Capability | Status | What it means |
|---|---|---|
| Visual Understanding | ✓ Supported | Image and document analysis support |
| Audio Processing | Not advertised | Speech and voice aligned flows |
| Tool Calling | Not advertised | Supports tools / function calling |
| Self-Hosting | ✓ Supported | Deploy outside managed APIs |
Capability signals
Directional signals derived from model metadata and capability tags — not official benchmark submissions.
Quick start
Call Qwen: Qwen2.5-VL 7B Instruct through an OpenAI-compatible client.
from openai import OpenAI
client = OpenAI(
base_url="https://openrouter.ai/api/v1",
api_key="YOUR_API_KEY"
)
response = client.chat.completions.create(
model="qwen/qwen-2.5-vl-7b-instruct",
messages=[{"role": "user", "content": "Explain quantum physics."}]
)
print(response.choices[0].message.content)Alternatives
Closest models by context window from a different provider.
| Model | Provider | Context | Input /1M | Actions |
|---|---|---|---|---|
| AionLabs: Aion-RP 1.0 (8B) | Aion Labs | 33K | $0.80 | View → |
| Arcee AI: Coder Large | Arcee Ai | 33K | $0.50 | View → |
Sources & attribution
Pricing and metadata are maintained in the ModelsAtlas catalog.
Frequently asked questions
How much does Qwen: Qwen2.5-VL 7B Instruct cost?
$0.20 per 1M input tokens and $0.20 per 1M output tokens on the official Qwen API.
What is the context window of Qwen: Qwen2.5-VL 7B Instruct?
Qwen: Qwen2.5-VL 7B Instruct supports up to 33K tokens of context.
Does Qwen: Qwen2.5-VL 7B Instruct support tool / function calling?
Qwen: Qwen2.5-VL 7B Instruct does not advertise first-class tool/function calling. Check vendor docs for the latest capabilities.
How do I access Qwen: Qwen2.5-VL 7B Instruct?
Use the official Qwen API with the model id `qwen/qwen-2.5-vl-7b-instruct`. See the Quick start section above for code examples.