Terminal-Bench leaderboard
Current leader: Claude Opus 4.7 at 90.2 (% resolved).
Cost vs quality
Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).
Closed model Open-weight model
Full leaderboard
#ModelScore% resolvedOutput $/1MReleased
1
90.2$150.00Apr 2026
2
84.7$30.00Apr 2026
3
81.8$15.00Mar 2026
4
80.2$12.00Feb 2026
5
79.8$25.00Feb 2026
6
78.4$14.00Feb 2026
7
69.4—Nov 2025
8
64.9—Dec 2025
9
64.3—Dec 2025
10
63.1—Nov 2025
11
57.3$2.50Feb 2026
12
53.4$15.00Feb 2026
13
OPEN
52.4$2.55Feb 2026
14
49.6—Aug 2025
15
47.6—Nov 2025
16
46.5—Sep 2025
17
OPEN
45.1—Mar 2026
18
OPEN
43.2$2.85Feb 2026
19
OPEN
42.7—Feb 2026
20
OPEN
39.6—Dec 2025
21
38—Aug 2025
22
OPEN
35.7—Nov 2025
23
35.5—Oct 2025
24
34.8—Aug 2025
25
OPEN
33.4—Dec 2025
26
32.6—Jun 2025
27
OPEN
27.8—Nov 2025
28
27.2—Sep 2025
29
24.6$1.00Apr 2026
30
OPEN
24.5—Sep 2025
31
21.8—Aug 2025
32
OPEN
18.7—Aug 2025
33
17.1—Sep 2025
34
17.1—Apr 2025
Scores from Epoch AI, updated —. Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.
