LiveBench leaderboard
LiveBench leaderboard: contamination-free LLM scores across reasoning, coding, agentic, math and more — plotted against measured cost per task so you see the best value. Current leader: Claude Fable 5 Max Effort at 82.97 (Overall (0–100)).
Cost vs quality
Each dot is a model. The dashed line is the value frontier — top-left is the best trade-off (high score, low cost).
Full leaderboard
Scores from LiveBench (2026-06-25), updated —. Cost is joined from our daily API pricing where a model matches; models we don't price still appear with their score. Always verify on the benchmark's own site before relying on a number.
What LiveBench measures
LiveBench is a general-capability benchmark built to resist the biggest problem with static leaderboards — training-data contamination. It draws questions from frequently-updated sources and rotates them over time, so models can’t simply have memorised the answers. Every question has an objective ground-truth answer, and scoring is fully automatic — there’s no LLM acting as judge, which removes a common source of bias.
Tasks span six areas, each scored and then combined into the 0–100 overall you see here:
- Reasoning — logic and multi-step problem solving.
- Coding — code generation and completion.
- Agentic coding — longer, tool-using software tasks.
- Mathematics — competition and applied math.
- Data analysis — tables, extraction, transformation.
- Language — comprehension and instruction following.
How to read this page
The leaderboard shows each model’s overall score plus its Coding and Agentic sub-scores — the two most relevant to developers. The cost-vs-quality chart uses LiveBench’s own measured cost per successful task on the horizontal axis, so you’re seeing real spend to get a task done, not just a headline token price. Reasoning-heavy models often score well but sit far to the right because they emit many tokens per task.
Limitations
LiveBench only covers models its maintainers have run, so very new or niche models may be missing. The overall score weights all categories equally — if you only care about coding, read the Coding column directly rather than the overall.
