Multimodal generalists
Multimodal generalists reason across text, images, audio, and video in a single prompt. Useful for universal assistants and content workflows.
9 models in this category. See also the “Best AI vision models” ranking with editorial picks and FAQ.
| Model | Provider | Context | Input price / 1M | Tier |
|---|---|---|---|---|
| Claude Opus 4.6 anthropic/claude-opus-4.6 | Anthropic | 1M Ultra context (1M+) | Custom | Variable |
| Claude Opus 4.7 anthropic/claude-opus-4.7 | Anthropic | 1M Ultra context (1M+) | Custom | Variable |
| Claude Sonnet 4.6 anthropic/claude-sonnet-4.6 | Anthropic | 1M Ultra context (1M+) | Custom | Variable |
| Claude Haiku 4.5 anthropic/claude-haiku-4.5 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Opus 4 anthropic/claude-opus-4 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Opus 4.1 anthropic/claude-opus-4.1 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Opus 4.5 anthropic/claude-opus-4.5 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Sonnet 4 anthropic/claude-sonnet-4 | Anthropic | 200K Long context (128K+) | Custom | Variable |
| Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 200K Long context (128K+) | Custom | Variable |