Use case

Best AI models for math

Math performance correlates strongly with overall reasoning. The leaders here are also the leaders on GPQA and MMLU.

Top 25 models for math problems

#ModelProviderContextInput price / 1MTier
1AionLabs: Aion-1.0-MiniAion Labs131K
Long context (128K+)
$0.70Budget
2Anthropic: Claude 3.7 SonnetAnthropic200K
Long context (128K+)
$3.00Standard
3Anthropic: Claude 3.7 Sonnet (thinking)Anthropic200K
Long context (128K+)
$3.00Standard
4Baidu: ERNIE 4.5 21B A3B ThinkingBaidu131K
Long context (128K+)
$0.07Budget
5DeepSeek: R1 Distill Qwen 32BDeepSeek33K
Short/standard context
$0.29Budget
6Google: Gemini 2.5 FlashGoogle1.0M
Ultra context (1M+)
$0.30Budget
7Google: Gemini 2.5 Flash (batch)Google1.0M
Ultra context (1M+)
$0.15Budget
8Google: Gemini 2.5 ProGoogle1.0M
Ultra context (1M+)
$1.25Standard
9Google: Gemini 2.5 Pro (batch)Google1.0M
Ultra context (1M+)
$0.63Budget
10Google: Gemini 2.5 Pro Preview 05-06Google1.0M
Ultra context (1M+)
$1.25Standard
11Google: Gemini 2.5 Pro Preview 06-05Google1.0M
Ultra context (1M+)
$1.25Standard
12Google: Gemini 3 Pro PreviewGoogle1.0M
Ultra context (1M+)
$2.00Standard
13Qwen: Qwen3 8BQwen131K
Long context (128K+)
$0.12Budget
14NVIDIA: Llama 3.3 Nemotron Super 49B V1.5Nvidia131K
Long context (128K+)
$0.10Budget
15NVIDIA: Nemotron Nano 12B 2 VLNvidia131K
Long context (128K+)
$0.20Budget
16OpenAI: o3OpenAI200K
Long context (128K+)
$2.00Standard
17OpenAI: o3 (batch)OpenAI200K
Long context (128K+)
$1.00Standard
18OpenAI: o3 MiniOpenAI200K
Long context (128K+)
$1.10Standard
19OpenAI: o3 Mini (batch)OpenAI200K
Long context (128K+)
$0.55Budget
20OpenAI: o3 Mini HighOpenAI200K
Long context (128K+)
$1.10Standard
21OpenAI: o3 Mini High (batch)OpenAI200K
Long context (128K+)
$0.55Budget
22Prime Intellect: INTELLECT-3Prime Intellect131K
Long context (128K+)
$0.20Budget
23Qwen: Qwen3 235B A22BQwen131K
Long context (128K+)
$0.46Budget
24Qwen: Qwen3 Next 80B A3B ThinkingQwen262K
Long context (128K+)
$0.15Budget
25Qwen: Qwen3 VL 235B A22B ThinkingQwen131K
Long context (128K+)
$0.40Budget

Frequently asked questions

What's the best model for math homework help?

Reasoning models (o-series, Claude thinking, DeepSeek R1) lead. For free options, DeepSeek R1 distilled variants run locally.

Related use cases