Category

Multimodal generalists

Multimodal generalists reason across text, images, audio, and video in a single prompt. Useful for universal assistants and content workflows.

9 models in this category. See also the Best AI vision models ranking with editorial picks and FAQ.

ModelProviderContextInput price / 1MTier
Claude Opus 4.6
anthropic/claude-opus-4.6
Anthropic1M
Ultra context (1M+)
CustomVariable
Claude Opus 4.7
anthropic/claude-opus-4.7
Anthropic1M
Ultra context (1M+)
CustomVariable
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6
Anthropic1M
Ultra context (1M+)
CustomVariable
Claude Haiku 4.5
anthropic/claude-haiku-4.5
Anthropic200K
Long context (128K+)
CustomVariable
Claude Opus 4
anthropic/claude-opus-4
Anthropic200K
Long context (128K+)
CustomVariable
Claude Opus 4.1
anthropic/claude-opus-4.1
Anthropic200K
Long context (128K+)
CustomVariable
Claude Opus 4.5
anthropic/claude-opus-4.5
Anthropic200K
Long context (128K+)
CustomVariable
Claude Sonnet 4
anthropic/claude-sonnet-4
Anthropic200K
Long context (128K+)
CustomVariable
Claude Sonnet 4.5
anthropic/claude-sonnet-4.5
Anthropic200K
Long context (128K+)
CustomVariable

Browse other categories