← All benchmarks·Source: Epoch AI

MATH (Level 5) leaderboard

Current leader: GPT-5 at 98.13 (% accuracy).

Cost vs quality

Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).

Not enough priced models on this benchmark yet to plot cost vs quality. The leaderboard below shows all scores.

Full leaderboard

#ModelScore% accuracyOutput $/1MReleased
1
GPT-5
98.13Aug 2025
2
GPT-5 mini
97.85Aug 2025
3
o4-mini
97.83Apr 2025
4
o3
97.77Dec 2024
5
Claude Sonnet 4.5
97.73Sep 2025
6
Qwen3-Max
97.13$3.90Jan 2026
7
DeepSeek-R1 (May 2025)OPEN
96.64May 2025
8
o3-mini
96.49Jan 2025
9
Claude Haiku 4.5
96.36Oct 2025
10
Gemini 2.5 Pro (May 2025)
95.9May 2025
11
Gemini 2.5 Pro (Mar 2025)
95.56Mar 2025
12
GPT-5 nano
95.24Aug 2025
13
o1
94.71Dec 2024
14
DeepSeek-R1OPEN
93.05Jan 2025
15
Claude 3.7 Sonnet
91.16Feb 2025
16
Grok-3 mini
90.94Feb 2025
17
o1-mini
89.18Sep 2024
18
Grok 3
88.75Feb 2025
19
GPT-4.1 mini
87.29Apr 2025
20
Claude Opus 4
85.05May 2025
21
Claude Sonnet 4
84.37May 2025
22
Gemini 2.0 Pro
83.46Dec 2024
23
GPT-4.1
83.01Apr 2025
24
Gemini 2.0 Flash (Feb 2025)
82.17Dec 2024
25
o1-preview
81.65Dec 2024
26
Mistral Medium 3
81.63May 2025
27
GPT-4.5
78.63Feb 2025
28
DeepSeek-V3 (Mar 2025)OPEN
75.55Mar 2025
29
Gemma 3 27BOPEN
74.04Mar 2025
30
Llama 4 MaverickOPEN
73.02Apr 2025
31
Gemini 1.5 Pro (Sept 2024)
70.39Sep 2024
32
GPT-4.1 nano
70Apr 2025
33
Qwen3-235B-A22BOPEN
68.86Apr 2025
34
Qwen2.5-Max
67.18Jan 2025
35
Phi-4OPEN
64.94Apr 2025
36
DeepSeek-V3OPEN
64.85Dec 2024
37
Grok-2 (Dec 2024)
63.52Aug 2024
38
Qwen2.5-72BOPEN
63.17Sep 2024
39
Llama 4 ScoutOPEN
62.27Apr 2025
40
Gemini 1.5 Flash (Sep 2024)
61.87May 2024
41
Claude 3.5 Sonnet (October 2024)
56.95Oct 2024
42
GPT-4o (Aug 2024)
53.28Aug 2024
43
GPT-4o mini
52.63Jul 2024
44
Claude 3.5 Sonnet
51.68Jun 2024
45
GPT-4o (May 2024)
51.05May 2024
46
Mistral Large 2 (Nov 2024)
50.28Jul 2024
47
Llama 3.1-405BOPEN
49.77Jul 2024
48
GPT-4o (Nov 2024)
49.77Nov 2024
49
Mistral Small 3.1OPEN
46.77Mar 2025
50
GPT-4 Turbo (Apr 2024)
46.73Apr 2024
51
Claude 3.5 Haiku
46.36Oct 2024
52
Mistral Large 2 (Jul 2024)
44.82Jul 2024
53
Llama 3.3 70BOPEN
41.6Dec 2024
54
Gemini 1.5 Pro (May 2024)
40.75May 2024
55
Llama 3.2 90BOPEN
39.44Sep 2024
56
Qwen2-72BOPEN
39.07Jun 2024
57
Claude 3 Opus
37.48Mar 2024
58
Llama 3.1-70BOPEN
36.68Jul 2024
59
Gemma 2 27BOPEN
27.89Jun 2024
60
Gemini 1.5 Flash (May 2024)
25.12May 2024
61
Mistral Large
24.46Feb 2024
62
Mixtral 8x22BOPEN
24.24Apr 2024
63
GPT-4 (Jun 2023)
22.97Jun 2023
64
Llama 3.1-8BOPEN
22.88Jul 2024
65
Llama 3-70BOPEN
22.55Apr 2024
66
Gemma 2 9BOPEN
21.01Jun 2024
67
Claude 3 Sonnet
18.17Mar 2024
68
phi-3-medium 14BOPEN
17.56Apr 2024
69
GPT-3.5 Turbo (Nov 2023)
15.89Jun 2023
70
Claude 3 Haiku
14.88Mar 2024
71
Claude 2
11.73Jul 2023
72
GPT-3.5 Turbo (Jan 2024)
11.63Jun 2023
73
Gemini 1.0 Pro
11.24Dec 2023
74
Mistral NeMoOPEN
10.83Jul 2024
75
Mixtral 8x7BOPEN
9.95Dec 2023
76
Llama 3-8BOPEN
6.13Apr 2024
77
Yi-34BOPEN
5.15Nov 2023
78
Llama 2-70BOPEN
3.29Jul 2023

Scores from Epoch AI, updated . Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.