Atlas / Models / Qwen / Qwen: Qwen2.5-VL 7B Instruct

Qwen: Qwen2.5-VL 7B Instruct✓ Catalog verified

qwen/qwen-2.5-vl-7b-instruct

Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements: - SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. - Understanding videos of 20min+: Qwen2.5-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. - Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2.5-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions. - Multilingual Support: to serve global users, besides English and Chinese, Qwen2.5-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc. For more details, see this blog post and GitHub repo. Usage of this model is subject to Tongyi Qianwen LICENSE AGREEMENT.

Input price
$0.20 /1M
Output price
$0.20 /1M
Context
33K
Modalities
textimage
Released
Aug 28, 2024
Tool calling
Atlas signal
78/100
01

Overview

Qwen2.5 VL 7B is a multimodal LLM from the Qwen Team with the following key enhancements: - SoTA understanding of images of various resolution & ratio: Qwen2.5-VL achieves state-of-the-art performance on visual understanding benchmarks, including MathVista, DocVQA, RealWorldQA, MTVQA, etc. - Understanding videos of 20min+: Qwen2.5-VL can understand videos over 20 minutes for high-quality video-based question answering, dialog, content creation, etc. - Agent that can operate your mobiles, robots, etc.: with the abilities of complex reasoning and decision making, Qwen2.5-VL can be integrated with devices like mobile phones, robots, etc., for automatic operation based on visual environment and text instructions. - Multilingual Support: to serve global users, besides English and Chinese, Qwen2.5-VL now supports the understanding of texts in different languages inside images, including most European languages, Japanese, Korean, Arabic, Vietnamese, etc. For more details, see this blog post and GitHub repo. Usage of this model is subject to Tongyi Qianwen LICENSE AGREEMENT.

Access: available through the official Qwen API. Context window: 33K.

Pricing and metadata from the ModelsAtlas catalog.Last refreshed Aug 5, 2026
02

Specifications

API identifierqwen/qwen-2.5-vl-7b-instruct
ProviderQwen
Model typeMultimodal LLM
Context window33K catalog
Input modalitiestext · image
Output modalitiestext
ReleasedAug 28, 2024
TokenizerQwen
ModeratedNo
Architecture modalitytext+image->text
Supported parameters
frequency_penaltylogit_biasmax_tokensmin_ppresence_penaltyrepetition_penaltyseedstoptemperaturetop_ktop_p
Source: provider documentation + catalog feedMethodology →
03

Pricing

Live pricing components for qwen/qwen-2.5-vl-7b-instruct as published in the catalog.

$0.20
Input /1M
$0.20
Output /1M
Per request
Per image
Pricing componentRaw unit priceNormalizedUnit
Prompt tokens$2e-7$0.20 / 1MPer input token
Completion tokens$2e-7$0.20 / 1MPer output token
Request feeN/AN/APer request
Image feeN/AN/APer image unit
Web search feeN/AN/APer search request
Source: catalog pricing feedCompare all pricing →Cheapest models →
04

Cost calculator

Estimate monthly spend from your own token volumes.

Scale / Volume

Estimated Total Cost

$0.14Calculated from current list pricing in the ModelsAtlas catalog.
05

Capabilities

CapabilityStatusWhat it means
Visual Understanding✓ SupportedImage and document analysis support
Audio ProcessingNot advertisedSpeech and voice aligned flows
Tool CallingNot advertisedSupports tools / function calling
Self-Hosting✓ SupportedDeploy outside managed APIs
Derived from catalog capability tags and supported parameters
06

Capability signals

Directional signals derived from model metadata and capability tags — not official benchmark submissions.

Qwen: Qwen2.5-VL 7B Instruct
MMLU Signal (Reasoning)
80signal
Coding Signal (HumanEval proxy)
74signal
Math Signal (GSM8K proxy)
83signal
Science Signal (GPQA proxy)
74signal
Estimated from metadata — not an official benchmark runAll benchmarks →Methodology →
07

Quick start

Call Qwen: Qwen2.5-VL 7B Instruct through an OpenAI-compatible client.

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_API_KEY"
)

response = client.chat.completions.create(
    model="qwen/qwen-2.5-vl-7b-instruct",
    messages=[{"role": "user", "content": "Explain quantum physics."}]
)

print(response.choices[0].message.content)
08

Alternatives

Closest models by context window from a different provider.

09

Sources & attribution

Pricing and metadata are maintained in the ModelsAtlas catalog.

Capability bars use metadata tags — directional estimates onlyHow we source and verify data →
10

Frequently asked questions

How much does Qwen: Qwen2.5-VL 7B Instruct cost?

$0.20 per 1M input tokens and $0.20 per 1M output tokens on the official Qwen API.

What is the context window of Qwen: Qwen2.5-VL 7B Instruct?

Qwen: Qwen2.5-VL 7B Instruct supports up to 33K tokens of context.

Does Qwen: Qwen2.5-VL 7B Instruct support tool / function calling?

Qwen: Qwen2.5-VL 7B Instruct does not advertise first-class tool/function calling. Check vendor docs for the latest capabilities.

How do I access Qwen: Qwen2.5-VL 7B Instruct?

Use the official Qwen API with the model id `qwen/qwen-2.5-vl-7b-instruct`. See the Quick start section above for code examples.