Vision models
Models that accept image inputs alongside text — for OCR, visual reasoning, screenshot analysis, and PDF understanding.
17 models in this category. See also the “Best AI vision models” ranking with editorial picks and FAQ.
| Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|
| Claude Opus 4.6 anthropic/claude-opus-4.6 | Anthropic | 1M Ultra context (1M+) | Custom | Variable |
| Claude Opus 4.7 anthropic/claude-opus-4.7 | Anthropic | 1M Ultra context (1M+) | Custom | Variable |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4.6 | Anthropic | 1M Ultra context (1M+) | Custom | Variable |
| Claude Haiku 4.5 anthropic/claude-haiku-4.5 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Opus 4 anthropic/claude-opus-4 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Opus 4.1 anthropic/claude-opus-4.1 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Opus 4.5 anthropic/claude-opus-4.5 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Sonnet 4 anthropic/claude-sonnet-4 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude 3 Opus anthropic/claude-3-opus | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude 3.5 Sonnet anthropic/claude-3.5-sonnet | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude 3.7 Sonnet anthropic/claude-3.7-sonnet | Anthropic | 200K Long context (128K+) | Custom | Variable |
| GPT-4o mini openai/gpt-4o-mini | OpenAI | 128K Long context (128K+) | Custom | Variable |
| Gemini 1.5 Flash google/gemini-1.5-flash | 1M Ultra context (1M+) | Custom | Variable | |
| Gemini 1.5 Pro google/gemini-1.5-pro | 1M Ultra context (1M+) | Custom | Variable | |
| Gemini 2.0 Flash google/gemini-2.0-flash | 1M Ultra context (1M+) | Custom | Variable | |
| GPT-4o openai/gpt-4o | OpenAI | 128K Long context (128K+) | Custom | Variable |