Category

Vision models

Models that accept image inputs alongside text — for OCR, visual reasoning, screenshot analysis, and PDF understanding.

17 models in this category. See also the Best AI vision models ranking with editorial picks and FAQ.

ModelProviderContextInput price / 1MTier
Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic1M
Ultra context (1M+)
CustomVariable
Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic1M
Ultra context (1M+)
CustomVariable
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic1M
Ultra context (1M+)
CustomVariable
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropic200K
Long context (128K+)
CustomVariable
Claude Opus 4
anthropic/claude-opus-4
Anthropic200K
Long context (128K+)
CustomVariable
Claude Opus 4.1
anthropic/claude-opus-4.1
Anthropic200K
Long context (128K+)
CustomVariable
Claude Opus 4.5
anthropic/claude-opus-4.5
Anthropic200K
Long context (128K+)
CustomVariable
Claude Sonnet 4
anthropic/claude-sonnet-4
Anthropic200K
Long context (128K+)
CustomVariable
Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic200K
Long context (128K+)
CustomVariable
Claude 3 Opus
anthropic/claude-3-opus
Anthropic200K
Long context (128K+)
CustomVariable
Claude 3.5 Sonnet
anthropic/claude-3.5-sonnet
Anthropic200K
Long context (128K+)
CustomVariable
Claude 3.7 Sonnet
anthropic/claude-3.7-sonnet
Anthropic200K
Long context (128K+)
CustomVariable
GPT-4o mini
openai/gpt-4o-mini
OpenAI128K
Long context (128K+)
CustomVariable
Gemini 1.5 Flash
google/gemini-1.5-flash
Google1M
Ultra context (1M+)
CustomVariable
Gemini 1.5 Pro
google/gemini-1.5-pro
Google1M
Ultra context (1M+)
CustomVariable
Gemini 2.0 Flash
google/gemini-2.0-flash
Google1M
Ultra context (1M+)
CustomVariable
GPT-4o
openai/gpt-4o
OpenAI128K
Long context (128K+)
CustomVariable

Browse other categories