← All benchmarks·Source: Epoch AI

Humanity's Last Exam leaderboard

Current leader: Gemini 3.1 Pro at 43.74 (% accuracy).

Cost vs quality

Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).

Closed model Open-weight model

Full leaderboard

#ModelScore% accuracyOutput $/1MReleased
1
Gemini 3.1 Pro
43.74$12.00Feb 2026
2
GPT-5.4 Pro
41.51$180.00Mar 2026
3
Muse Spark
37.56Apr 2026
4
Gemini 3 Pro
34.37Nov 2025
5
GPT-5.4
33.03$15.00Mar 2026
6
Claude Opus 4.7
32.98$150.00Apr 2026
7
Claude Opus 4.6
31.13$25.00Feb 2026
8
GPT-5 Pro
28.19Oct 2025
9
GPT-5.2
24.16Dec 2025
10
GPT-5
21.55Aug 2025
11
Claude Opus 4.5
21.43Nov 2025
12
Kimi K2.5OPEN
20.56$2.85Feb 2026
13
GPT-5.1
19.83Nov 2025
14
Gemini 2.5 Pro (Jun 2025)
17.69Jun 2025
15
o3
16.3Dec 2024
16
GPT-5 mini
15.38Aug 2025
17
Gemini 2.5 Pro (Mar 2025)
14.03Mar 2025
18
o4-mini
13.95Apr 2025
19
Gemini 2.5 Pro (May 2025)
13.66May 2025
20
Claude Sonnet 4.5
9.37Sep 2025
21
Gemini 2.5 Flash (Apr 2025)
7.65Apr 2025
22
Claude Opus 4.1
7.06Aug 2025
23
Gemini 2.5 Flash (May 2025)
6.47May 2025
24
Claude Opus 4
6.22May 2025
25
Gemini 3.1 Flash-Lite
4.03$1.50Mar 2026
26
Claude 3.7 Sonnet
3.4Feb 2025
27
o1
3.32Dec 2024
28
Claude Sonnet 4
3.11May 2025
29
Gemini 2.0 Flash Thinking (Jan 2025)
1.85Jan 2025
30
Llama 4 MaverickOPEN
0.92Apr 2025
31
GPT-4.5
0.67Feb 2025
32
GPT-4.1
0.63Apr 2025
33
Amazon Nova Pro
0Dec 2024
34
Claude 3.5 Sonnet (October 2024)
0Oct 2024
35
Mistral Medium 3
0May 2025
36
GPT-4o (Nov 2024)
0Nov 2024
37
Gemini 1.5 Pro (Sept 2024)
0Sep 2024

Scores from Epoch AI, updated . Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.