Humanity's Last Exam leaderboard
Current leader: Gemini 3.1 Pro at 43.74 (% accuracy).
Cost vs quality
Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).
Closed model Open-weight model
Full leaderboard
#ModelScore% accuracyOutput $/1MReleased
1
43.74$12.00Feb 2026
2
41.51$180.00Mar 2026
3
37.56—Apr 2026
4
34.37—Nov 2025
5
33.03$15.00Mar 2026
6
32.98$150.00Apr 2026
7
31.13$25.00Feb 2026
8
28.19—Oct 2025
9
24.16—Dec 2025
10
21.55—Aug 2025
11
21.43—Nov 2025
12
OPEN
20.56$2.85Feb 2026
13
19.83—Nov 2025
14
17.69—Jun 2025
15
16.3—Dec 2024
16
15.38—Aug 2025
17
14.03—Mar 2025
18
13.95—Apr 2025
19
13.66—May 2025
20
9.37—Sep 2025
21
7.65—Apr 2025
22
7.06—Aug 2025
23
6.47—May 2025
24
6.22—May 2025
25
4.03$1.50Mar 2026
26
3.4—Feb 2025
27
3.32—Dec 2024
28
3.11—May 2025
29
1.85—Jan 2025
30
OPEN
0.92—Apr 2025
31
0.67—Feb 2025
32
0.63—Apr 2025
33
0—Dec 2024
34
0—Oct 2024
35
0—May 2025
36
0—Nov 2024
37
0—Sep 2024
Scores from Epoch AI, updated —. Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.
