Which model is the best value, not just the top score.
11 independent benchmarks — coding, reasoning, math and general knowledge — each plotted against price. Pick a benchmark and read cost vs quality directly.
LiveBench — cost vs quality
Each dot is a model; the dashed line is the value frontier. Top-left is best (high score, low cost).
Source: LiveBench · Overall (0–100). Cost is joined from our daily pricing where a model matches; models we don't price still appear (score only).
Benchmark data updated Jul 2026 from LiveBench (livebench.ai) and Epoch AI (epoch.ai). Scores belong to their respective benchmarks; we join our daily pricing to show cost vs quality.
How to read these benchmarks
A leaderboard tells you which model scores highest. It doesn’t tell you which model is the smart buy — the top model can cost many times more than one a point or two behind it. That’s why every benchmark here is plotted as cost vs quality: score on the vertical axis, price or measured cost on the horizontal. Models toward the top-left give you the most capability per dollar, and the dashed value frontier traces the best trade-off at each price point.
We started with three widely-cited coding-focused benchmarks, and keep adding more as Epoch AI's dataset covers them — each measures something different:
- LiveBench — a broad, regularly-refreshed benchmark spanning reasoning, coding, agentic tasks, math, data analysis and language, scored 0–100. Its questions rotate to limit training-data contamination, and it publishes a measured cost-per-task.
- SWE-bench Verified — a human-validated set of real GitHub issues; the score is the percentage a model resolves with a working code patch. The closest thing to “can it actually fix bugs in a real repo.”
- Aider Polyglot — code-editing exercises across multiple programming languages; the score is the percentage solved with correctly-formatted edits. A good read on day-to-day coding-assistant reliability.
Beyond coding, we also track general knowledge (MMLU), graduate-level science reasoning (GPQA Diamond), competition mathematics (MATH Level 5, AIME), extremely hard cross-domain questions (Humanity's Last Exam), abstract reasoning (ARC-AGI-2), classic multi-step reasoning (BBH), and agentic terminal use (Terminal-Bench) — all sourced from the same Epoch AI dataset, so the pipeline can keep adding more over time without new integrations. Every model on every benchmark is also automatically tagged as open-weight or closed — see the open-vs-closed race to compare the two groups directly.
No single benchmark is the whole truth — they test different skills and each has blind spots. Read them together, and weigh the scores against the cost. Scores come straight from each benchmark’s published data; see methodology for sources and refresh cadence.
