Big-Bench Hard leaderboard
Current leader: Gemini 1.5 Pro (May 2024) at 85.6 (% accuracy).
Cost vs quality
Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).
Not enough priced models on this benchmark yet to plot cost vs quality. The leaderboard below shows all scores.
Full leaderboard
#ModelScore% accuracyOutput $/1MReleased
1
85.6—May 2024
2
OPEN
83.33—Dec 2024
3
OPEN
77.2—Jul 2024
4
OPEN
75.2—Apr 2024
5
OPEN
73.07—Sep 2024
6
OPEN
72.13—Apr 2024
7
71.73—May 2024
8
66.83—Jun 2023
9
OPEN
62.27—Nov 2023
10
OPEN
62.27—Apr 2024
11
OPEN
59.07—Jul 2023
12
OPEN
53.2—Jul 2023
13
48.79—Jun 2023
14
OPEN
45.87—Dec 2023
15
44.93—Feb 2024
16
OPEN
44.53—Feb 2023
17
OPEN
44.27—Jul 2023
18
41.47—Oct 2023
19
OPEN
40.13—Feb 2024
20
OPEN
40—Sep 2023
21
36.67—Sep 2023
22
OPEN
33.33—Feb 2023
23
OPEN
32—Sep 2023
24
OPEN
29.6—Nov 2023
25
OPEN
26.67—Sep 2023
26
25.47—Jul 2023
27
24.05—Apr 2023
28
OPEN
22.13—Sep 2023
29
OPEN
18.88—Jul 2023
30
OPEN
17.33—Jun 2023
31
OPEN
17.2—Feb 2023
32
OPEN
16.13—Mar 2023
33
16—Jul 2023
34
OPEN
14.13—May 2023
35
OPEN
13.6—Feb 2024
36
13.13—Nov 2024
37
11.6—Jun 2023
38
OPEN
11.33—Feb 2023
39
OPEN
9.97—Jun 2023
40
OPEN
5.03—Apr 2023
41
4.27—Nov 2023
Scores from Epoch AI, updated —. Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.
