← All benchmarks·Source: Epoch AI

Terminal-Bench leaderboard

Current leader: Claude Opus 4.7 at 90.2 (% resolved).

Cost vs quality

Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).

Closed model Open-weight model

Full leaderboard

#ModelScore% resolvedOutput $/1MReleased
1
Claude Opus 4.7
90.2$150.00Apr 2026
2
GPT-5.5
84.7$30.00Apr 2026
3
GPT-5.4
81.8$15.00Mar 2026
4
Gemini 3.1 Pro
80.2$12.00Feb 2026
5
Claude Opus 4.6
79.8$25.00Feb 2026
6
GPT-5.3 Codex
78.4$14.00Feb 2026
7
Gemini 3 Pro
69.4Nov 2025
8
GPT-5.2
64.9Dec 2025
9
Gemini 3 Flash
64.3Dec 2025
10
Claude Opus 4.5
63.1Nov 2025
11
Grok 4.20
57.3$2.50Feb 2026
12
Claude Sonnet 4.6
53.4$15.00Feb 2026
13
GLM-5OPEN
52.4$2.55Feb 2026
14
GPT-5
49.6Aug 2025
15
GPT-5.1
47.6Nov 2025
16
Claude Sonnet 4.5
46.5Sep 2025
17
MiniMax-M2.7OPEN
45.1Mar 2026
18
Kimi K2.5OPEN
43.2$2.85Feb 2026
19
MiniMax-M2.5OPEN
42.7Feb 2026
20
DeepSeek-V3.2OPEN
39.6Dec 2025
21
Claude Opus 4.1
38Aug 2025
22
Kimi K2 ThinkingOPEN
35.7Nov 2025
23
Claude Haiku 4.5
35.5Oct 2025
24
GPT-5 mini
34.8Aug 2025
25
GLM-4.7OPEN
33.4Dec 2025
26
Gemini 2.5 Pro (Jun 2025)
32.6Jun 2025
27
Kimi K2 (Jul 2025)OPEN
27.8Nov 2025
28
Grok 4
27.2Sep 2025
29
Qwen 3.6 35B-A3B
24.6$1.00Apr 2026
30
GLM-4.6OPEN
24.5Sep 2025
31
GPT-5 nano
21.8Aug 2025
32
gpt-oss-120bOPEN
18.7Aug 2025
33
Gemini 2.5 Flash (Sep 2025)
17.1Sep 2025
34
Gemini 2.5 Flash (Jun 2025)
17.1Apr 2025

Scores from Epoch AI, updated . Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.