Benchmarks

Which model is the best value, not just the top score.

11 independent benchmarks — coding, reasoning, math and general knowledge — each plotted against price. Pick a benchmark and read cost vs quality directly.

New
Open-weight vs closed models — the race, charted
See exactly how many months behind the closed-model frontier open-weight models (Llama, DeepSeek, Qwen, GLM, Kimi…) currently sit, per benchmark.
View the race →

LiveBench — cost vs quality

Each dot is a model; the dashed line is the value frontier. Top-left is best (high score, low cost).

Full LiveBench page →
Closed model Open-weight model
#ModelScoreOverallCodingAgentic$ / taskReleased
1
Claude Fable 5 Max Effort
82.9785.9962.17$1.57Jun 2026
2
GPT-5.6 Sol Max Effort
81.0583.9456.21$0.589Jul 2026
3
GPT-5.5 Thinking xHigh Effort
80.1982.1553.99$0.530Apr 2026
4
claude-opus-5-max-effort
80.0981.4565.2$0.641Jul 2026
5
kimi-k3OPEN
79.1981.4562.17$0.379Jul 2026
6
GPT-5.4 Thinking xHigh Effort
77.9777.5453.84$0.387Mar 2026
7
GPT-5.6 Terra Max Effort
77.9478.2554.95$0.497Jul 2026
8
Gemini 3.1 Pro Preview High
76.9576.4544.14$0.262Feb 2026
9
Claude 4.7 Opus Thinking xHigh Effort
76.5382.0950.66$0.528
10
claude-opus-4-8-max-effort
76.2281.8350.5$0.986
11
Claude Sonnet 5 xHigh Effort
76.0480.6859.39$0.492Jun 2026
12
Grok 4.5
75.7768.5956.46$0.128Jul 2026
13
muse-spark-1.1-xhigh
75.377.1658.54$0.234Jul 2026
14
Gemini 3.5 Flash High
74.6478.1848.99$0.249
15
GPT-5.2 High
74.6376.0750.25$0.234Dec 2025
16
Claude 4.6 Opus Thinking High Effort
74.5278.1848.99$0.404
17
GPT-5.2 Codex
73.9783.6249.39$0.187Dec 2025
18
gemini-3.6-flash-high
73.5977.8643.43$0.235Jul 2026
19
GPT-5.6 Luna Max Effort
73.5682.9148.43$0.202Jul 2026
20
GLM-5.2OPEN
73.1679.6551.77$0.196Jun 2026
21
Qwen 3.7 Max
73.1474.2243.59$0.182May 2026
22
Claude 4.6 Sonnet Thinking Medium Effort
72.9979.2742.63$0.306
23
Claude 4.5 Opus Thinking High Effort
72.5879.6539.7$0.610
24
inkling-xhighOPEN
71.9271.0249.39$0.343Jul 2026
25
DeepSeek V4 ProOPEN
71.5769.9942.63$0.050Apr 2026
26
Kimi K2.6 ThinkingOPEN
70.5478.5746.92$0.169Apr 2026
27
GPT-5.4 Nano xHigh
69.5870.8446.77$0.091Mar 2026
28
Qwen 3.6 Plus
68.978.1841.36$0.227Apr 2026
29
Kimi K2.7 CodeOPEN
68.4173.9645.66$0.100Jun 2026
30
Grok Build 0.1
67.7865.3945.81$0.024
31
Minimax M3
67.2668.240.66$0.060Jun 2026
32
GPT-5.4 Mini xHigh
66.3771.6241.67$0.334Mar 2026
33
DeepSeek V4 FlashOPEN
65.4869.2337.63$0.016Apr 2026
34
Qwen 3.6 27BOPEN
64.0371.7839.29$0.202Apr 2026
35
gemini-3.5-flash-lite-high
63.9476.0745.25$0.069Jul 2026
36
Grok 4.3
62.2569.9318.54$0.061

Source: LiveBench · Overall (0–100). Cost is joined from our daily pricing where a model matches; models we don't price still appear (score only).

Benchmark data updated Jul 2026 from LiveBench (livebench.ai) and Epoch AI (epoch.ai). Scores belong to their respective benchmarks; we join our daily pricing to show cost vs quality.

How to read these benchmarks

A leaderboard tells you which model scores highest. It doesn’t tell you which model is the smart buy — the top model can cost many times more than one a point or two behind it. That’s why every benchmark here is plotted as cost vs quality: score on the vertical axis, price or measured cost on the horizontal. Models toward the top-left give you the most capability per dollar, and the dashed value frontier traces the best trade-off at each price point.

We started with three widely-cited coding-focused benchmarks, and keep adding more as Epoch AI's dataset covers them — each measures something different:

  • LiveBench — a broad, regularly-refreshed benchmark spanning reasoning, coding, agentic tasks, math, data analysis and language, scored 0–100. Its questions rotate to limit training-data contamination, and it publishes a measured cost-per-task.
  • SWE-bench Verified — a human-validated set of real GitHub issues; the score is the percentage a model resolves with a working code patch. The closest thing to “can it actually fix bugs in a real repo.”
  • Aider Polyglot — code-editing exercises across multiple programming languages; the score is the percentage solved with correctly-formatted edits. A good read on day-to-day coding-assistant reliability.

Beyond coding, we also track general knowledge (MMLU), graduate-level science reasoning (GPQA Diamond), competition mathematics (MATH Level 5, AIME), extremely hard cross-domain questions (Humanity's Last Exam), abstract reasoning (ARC-AGI-2), classic multi-step reasoning (BBH), and agentic terminal use (Terminal-Bench) — all sourced from the same Epoch AI dataset, so the pipeline can keep adding more over time without new integrations. Every model on every benchmark is also automatically tagged as open-weight or closed — see the open-vs-closed race to compare the two groups directly.

No single benchmark is the whole truth — they test different skills and each has blind spots. Read them together, and weigh the scores against the cost. Scores come straight from each benchmark’s published data; see methodology for sources and refresh cadence.

Frequently asked questions

What's the difference between LiveBench, SWE-bench Verified and Aider Polyglot?
LiveBench is broad and multi-category (reasoning, coding, math, etc.), scored 0–100 with rotating questions to limit contamination. SWE-bench Verified measures the percentage of real, human-validated GitHub issues a model resolves with a working patch. Aider Polyglot measures the percentage of multi-language code-editing exercises solved with correctly formatted edits. They test different skills, so read them together.
What does 'cost vs quality' mean on the charts?
Each dot is a model plotted with its benchmark score on the vertical axis and its cost (measured cost-per-task for LiveBench, or output price per 1M tokens for the others) on the horizontal. Top-left is best — high score, low cost. The dashed line is the value frontier: the best score available at each price point.
Why do some models show a score but no cost?
We only show cost for models we currently price. Benchmarks include many models — some older, some we don’t track pricing for — so those appear on the leaderboard with a score but are omitted from the cost chart. Nothing is hidden; the score is still shown.
How current are the benchmark scores?
They are refreshed from each benchmark’s published data on a daily automated schedule and rebuilt into the site. The “updated” date under the tables shows the latest refresh. Always verify a specific number on the benchmark’s own site before relying on it.
Can I trust a single benchmark to pick a model?
No benchmark is definitive. Each measures a slice of capability and has blind spots or possible contamination. Use them as directional signals, cross-check across benchmarks, and validate on your own tasks before committing.
How do you tag a model as open-weight?
Automatically, from two independent data signals — whether a provider's API metadata links the model to a public Hugging Face repository, and Epoch AI's own model-accessibility classification — refreshed on our daily pipeline. There is no hand-maintained list, so new open-weight releases get tagged as soon as they appear in either source.