Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
Model details →Mistral: Mistral Large 3 2512 vs Pareto
Mistral: Mistral Large 3 2512 wins on 2 of 6 axes — pricing and capability skew in its favour for most workloads.
Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a broad range of general-purpose tasks.
Model details →Side-by-side comparison
| Capability | Mistral: Mistral Large 3 2512 | Pareto | Winner |
|---|---|---|---|
| Context window Maximum number of input tokens the model can attend to in a single request. | 262K | 262K | Tie |
| Input price (per 1M) Cost per million input tokens billed by the provider. | $0.50 | $2.50 | 🏆 Mistral: Mistral Large 3 2512 |
| Output price (per 1M) Cost per million output tokens billed by the provider. | $1.50 | $7.50 | 🏆 Mistral: Mistral Large 3 2512 |
| Tool / function calling First-class support for emitting structured tool calls. | Yes | Yes | Tie |
| Vision input Accepts image inputs alongside text. | Yes | Yes | Tie |
| Reasoning mode Internal chain-of-thought / extended-thinking support. | — | — | Tie |
Frequently asked questions
Is Mistral: Mistral Large 3 2512 better than Pareto?
Mistral: Mistral Large 3 2512 wins on 2 of 6 axes — pricing and capability skew in its favour for most workloads.
What's the price difference between Mistral: Mistral Large 3 2512 and Pareto?
Input: $0.50 vs $2.50 per 1M tokens. Output: $1.50 vs $7.50 per 1M tokens.
What context windows do Mistral: Mistral Large 3 2512 and Pareto support?
Mistral: Mistral Large 3 2512 supports up to 262K tokens. Pareto supports up to 262K tokens.
Do both Mistral: Mistral Large 3 2512 and Pareto support tool calling?
Mistral: Mistral Large 3 2512: yes. Pareto: yes.