Atlas / Models / Meta / Meta: Llama 3.2 11B Vision Instruct

Meta: Llama 3.2 11B Vision Instruct✓ Catalog verified

meta-llama/llama-3.2-11b-vision-instruct

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answering, bridging the gap between language generation and visual reasoning. Pre-trained on a massive dataset of image-text pairs, it performs well in complex, high-accuracy image analysis. Its ability to integrate visual understanding with language processing makes it an ideal solution for industries requiring comprehensive visual-linguistic AI applications, such as content creation, AI-driven customer service, and research. Click here for the original model card. Usage of this model is subject to Meta's Acceptable Use Policy.

Input price
$0.05 /1M
Output price
$0.05 /1M
Context
131K
Modalities
textimage
Released
Sep 25, 2024
Tool calling
Atlas signal
83/100
01

Overview

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and visual question answering, bridging the gap between language generation and visual reasoning. Pre-trained on a massive dataset of image-text pairs, it performs well in complex, high-accuracy image analysis. Its ability to integrate visual understanding with language processing makes it an ideal solution for industries requiring comprehensive visual-linguistic AI applications, such as content creation, AI-driven customer service, and research. Click here for the original model card. Usage of this model is subject to Meta's Acceptable Use Policy.

Access: available through the official Meta API. Context window: 131K.

Pricing and metadata from the ModelsAtlas catalog.Last refreshed Aug 5, 2026
02

Specifications

API identifiermeta-llama/llama-3.2-11b-vision-instruct
ProviderMeta
Model typeMultimodal LLM
Context window131K catalog
Input modalitiestext · image
Output modalitiestext
ReleasedSep 25, 2024
TokenizerLlama3
Instruction formatllama3
ModeratedNo
Architecture modalitytext+image->text
Supported parameters
frequency_penaltymax_tokensmin_ppresence_penaltyrepetition_penaltyresponse_formatseedstoptemperaturetop_ktop_p
Source: provider documentation + catalog feedMethodology →
03

Pricing

Live pricing components for meta-llama/llama-3.2-11b-vision-instruct as published in the catalog.

$0.05
Input /1M
$0.05
Output /1M
Per request
Per image
Pricing componentRaw unit priceNormalizedUnit
Prompt tokens$4.9e-8$0.05 / 1MPer input token
Completion tokens$4.9e-8$0.05 / 1MPer output token
Request feeN/AN/APer request
Image feeN/AN/APer image unit
Web search feeN/AN/APer search request
Source: catalog pricing feedCompare all pricing →Cheapest models →
04

Cost calculator

Estimate monthly spend from your own token volumes.

Scale / Volume

Estimated Total Cost

$0.03Calculated from current list pricing in the ModelsAtlas catalog.
05

Capabilities

CapabilityStatusWhat it means
Visual Understanding✓ SupportedImage and document analysis support
Audio ProcessingNot advertisedSpeech and voice aligned flows
Tool CallingNot advertisedSupports tools / function calling
Self-Hosting✓ SupportedDeploy outside managed APIs
Derived from catalog capability tags and supported parameters
06

Capability signals

Directional signals derived from model metadata and capability tags — not official benchmark submissions.

Meta: Llama 3.2 11B Vision Instruct
MMLU Signal (Reasoning)
85signal
Coding Signal (HumanEval proxy)
82signal
Math Signal (GSM8K proxy)
79signal
Science Signal (GPQA proxy)
87signal
Estimated from metadata — not an official benchmark runAll benchmarks →Methodology →
07

Quick start

Call Meta: Llama 3.2 11B Vision Instruct through an OpenAI-compatible client.

from openai import OpenAI

client = OpenAI(
    base_url="https://openrouter.ai/api/v1",
    api_key="YOUR_API_KEY"
)

response = client.chat.completions.create(
    model="meta-llama/llama-3.2-11b-vision-instruct",
    messages=[{"role": "user", "content": "Explain quantum physics."}]
)

print(response.choices[0].message.content)
08

Alternatives

Closest models by context window from a different provider.

09

Sources & attribution

Pricing and metadata are maintained in the ModelsAtlas catalog.

Capability bars use metadata tags — directional estimates onlyHow we source and verify data →
10

Frequently asked questions

How much does Meta: Llama 3.2 11B Vision Instruct cost?

$0.05 per 1M input tokens and $0.05 per 1M output tokens on the official Meta API.

What is the context window of Meta: Llama 3.2 11B Vision Instruct?

Meta: Llama 3.2 11B Vision Instruct supports up to 131K tokens of context.

Does Meta: Llama 3.2 11B Vision Instruct support tool / function calling?

Meta: Llama 3.2 11B Vision Instruct does not advertise first-class tool/function calling. Check vendor docs for the latest capabilities.

How do I access Meta: Llama 3.2 11B Vision Instruct?

Use the official Meta API with the model id `meta-llama/llama-3.2-11b-vision-instruct`. See the Quick start section above for code examples.