← All benchmarks·Source: LiveBench

LiveBench leaderboard

LiveBench leaderboard: contamination-free LLM scores across reasoning, coding, agentic, math and more — plotted against measured cost per task so you see the best value. Current leader: Claude Fable 5 Max Effort at 82.97 (Overall (0–100)).

Cost vs quality

Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).

Closed model Open-weight model

Full leaderboard

#ModelScoreOverallCodingAgentic$ / taskReleased
1
Claude Fable 5 Max Effort
82.9785.9962.17$1.57Jun 2026
2
GPT-5.6 Sol Max Effort
81.0583.9456.21$0.589Jul 2026
3
GPT-5.5 Thinking xHigh Effort
80.1982.1553.99$0.530Apr 2026
4
claude-opus-5-max-effort
80.0981.4565.2$0.641Jul 2026
5
kimi-k3OPEN
79.1981.4562.17$0.379Jul 2026
6
GPT-5.4 Thinking xHigh Effort
77.9777.5453.84$0.387Mar 2026
7
GPT-5.6 Terra Max Effort
77.9478.2554.95$0.497Jul 2026
8
Gemini 3.1 Pro Preview High
76.9576.4544.14$0.262Feb 2026
9
Claude 4.7 Opus Thinking xHigh Effort
76.5382.0950.66$0.528
10
claude-opus-4-8-max-effort
76.2281.8350.5$0.986
11
Claude Sonnet 5 xHigh Effort
76.0480.6859.39$0.492Jun 2026
12
Grok 4.5
75.7768.5956.46$0.128Jul 2026
13
muse-spark-1.1-xhigh
75.377.1658.54$0.234Jul 2026
14
Gemini 3.5 Flash High
74.6478.1848.99$0.249
15
GPT-5.2 High
74.6376.0750.25$0.234Dec 2025
16
Claude 4.6 Opus Thinking High Effort
74.5278.1848.99$0.404
17
GPT-5.2 Codex
73.9783.6249.39$0.187Dec 2025
18
gemini-3.6-flash-high
73.5977.8643.43$0.235Jul 2026
19
GPT-5.6 Luna Max Effort
73.5682.9148.43$0.202Jul 2026
20
GLM-5.2OPEN
73.1679.6551.77$0.196Jun 2026
21
Qwen 3.7 Max
73.1474.2243.59$0.182May 2026
22
Claude 4.6 Sonnet Thinking Medium Effort
72.9979.2742.63$0.306
23
Claude 4.5 Opus Thinking High Effort
72.5879.6539.7$0.610
24
inkling-xhighOPEN
71.9271.0249.39$0.343Jul 2026
25
DeepSeek V4 ProOPEN
71.5769.9942.63$0.050Apr 2026
26
Kimi K2.6 ThinkingOPEN
70.5478.5746.92$0.169Apr 2026
27
GPT-5.4 Nano xHigh
69.5870.8446.77$0.091Mar 2026
28
Qwen 3.6 Plus
68.978.1841.36$0.227Apr 2026
29
Kimi K2.7 CodeOPEN
68.4173.9645.66$0.100Jun 2026
30
Grok Build 0.1
67.7865.3945.81$0.024
31
Minimax M3
67.2668.240.66$0.060Jun 2026
32
GPT-5.4 Mini xHigh
66.3771.6241.67$0.334Mar 2026
33
DeepSeek V4 FlashOPEN
65.4869.2337.63$0.016Apr 2026
34
Qwen 3.6 27BOPEN
64.0371.7839.29$0.202Apr 2026
35
gemini-3.5-flash-lite-high
63.9476.0745.25$0.069Jul 2026
36
Grok 4.3
62.2569.9318.54$0.061

Scores from LiveBench (2026-06-25), updated . Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.

What LiveBench measures

LiveBench is a general-capability benchmark built to resist the biggest problem with static leaderboards — training-data contamination. It draws questions from frequently-updated sources and rotates them over time, so models can’t simply have memorised the answers. Every question has an objective ground-truth answer, and scoring is fully automatic — there’s no LLM acting as judge, which removes a common source of bias.

Tasks span six areas, each scored and then combined into the 0–100 overall you see here:

  • Reasoning — logic and multi-step problem solving.
  • Coding — code generation and completion.
  • Agentic coding — longer, tool-using software tasks.
  • Mathematics — competition and applied math.
  • Data analysis — tables, extraction, transformation.
  • Language — comprehension and instruction following.

How to read this page

The leaderboard shows each model’s overall score plus its Coding and Agentic sub-scores — the two most relevant to developers. The cost-vs-quality chart uses LiveBench’s own measured cost per successful task on the horizontal axis, so you’re seeing real spend to get a task done, not just a headline token price. Reasoning-heavy models often score well but sit far to the right because they emit many tokens per task.

Limitations

LiveBench only covers models its maintainers have run, so very new or niche models may be missing. The overall score weights all categories equally — if you only care about coding, read the Coding column directly rather than the overall.

Frequently asked questions

Is LiveBench contamination-free?
It is designed to strongly limit contamination by using frequently-updated questions that rotate over time, rather than a fixed public test set. No benchmark can guarantee zero contamination, but this design makes memorisation much harder than with static benchmarks.
How is the overall score calculated?
Each of the six categories (reasoning, coding, agentic coding, mathematics, data analysis, language) is scored, and the overall is their average on a 0–100 scale. TokenCost shows the overall plus the coding and agentic sub-scores.
What is the cost axis on the chart?
LiveBench’s own measured cost per successful task, in US dollars — the real spend to complete a task, which captures how many tokens a model uses, not just its per-token price.