← All benchmarks

Open-weight vs closed models: the race, charted.

Open-weight releases — Llama, DeepSeek, Qwen, GLM, Kimi, Gemma and others — are closing in on closed, API-only frontier models. This page plots both groups by release date and benchmark score, and computes exactly how many months behind the closed frontier the best open-weight model currently sits, per benchmark.

LiveBench — open-weight vs closed, over time

Each dot is a model release, plotted by its release date and score. The stepped lines trace the running-best ("frontier") for each group — 8 open-weight, 28 closed models tracked.

2.8months — how far behind the closed-model frontier the best open-weight model on LiveBench currently sits.
Show
#ModelScoreOverallCodingAgentic$ / taskReleased
1
Claude Fable 5 Max Effort
82.9785.9962.17$1.57Jun 2026
2
GPT-5.6 Sol Max Effort
81.0583.9456.21$0.589Jul 2026
3
GPT-5.5 Thinking xHigh Effort
80.1982.1553.99$0.530Apr 2026
4
claude-opus-5-max-effort
80.0981.4565.2$0.641Jul 2026
5
kimi-k3OPEN
79.1981.4562.17$0.379Jul 2026
6
GPT-5.4 Thinking xHigh Effort
77.9777.5453.84$0.387Mar 2026
7
GPT-5.6 Terra Max Effort
77.9478.2554.95$0.497Jul 2026
8
Gemini 3.1 Pro Preview High
76.9576.4544.14$0.262Feb 2026
9
Claude 4.7 Opus Thinking xHigh Effort
76.5382.0950.66$0.528
10
claude-opus-4-8-max-effort
76.2281.8350.5$0.986
11
Claude Sonnet 5 xHigh Effort
76.0480.6859.39$0.492Jun 2026
12
Grok 4.5
75.7768.5956.46$0.128Jul 2026
13
muse-spark-1.1-xhigh
75.377.1658.54$0.234Jul 2026
14
Gemini 3.5 Flash High
74.6478.1848.99$0.249
15
GPT-5.2 High
74.6376.0750.25$0.234Dec 2025
16
Claude 4.6 Opus Thinking High Effort
74.5278.1848.99$0.404
17
GPT-5.2 Codex
73.9783.6249.39$0.187Dec 2025
18
gemini-3.6-flash-high
73.5977.8643.43$0.235Jul 2026
19
GPT-5.6 Luna Max Effort
73.5682.9148.43$0.202Jul 2026
20
GLM-5.2OPEN
73.1679.6551.77$0.196Jun 2026
21
Qwen 3.7 Max
73.1474.2243.59$0.182May 2026
22
Claude 4.6 Sonnet Thinking Medium Effort
72.9979.2742.63$0.306
23
Claude 4.5 Opus Thinking High Effort
72.5879.6539.7$0.610
24
inkling-xhighOPEN
71.9271.0249.39$0.343Jul 2026
25
DeepSeek V4 ProOPEN
71.5769.9942.63$0.050Apr 2026
26
Kimi K2.6 ThinkingOPEN
70.5478.5746.92$0.169Apr 2026
27
GPT-5.4 Nano xHigh
69.5870.8446.77$0.091Mar 2026
28
Qwen 3.6 Plus
68.978.1841.36$0.227Apr 2026
29
Kimi K2.7 CodeOPEN
68.4173.9645.66$0.100Jun 2026
30
Grok Build 0.1
67.7865.3945.81$0.024
31
Minimax M3
67.2668.240.66$0.060Jun 2026
32
GPT-5.4 Mini xHigh
66.3771.6241.67$0.334Mar 2026
33
DeepSeek V4 FlashOPEN
65.4869.2337.63$0.016Apr 2026
34
Qwen 3.6 27BOPEN
64.0371.7839.29$0.202Apr 2026
35
gemini-3.5-flash-lite-high
63.9476.0745.25$0.069Jul 2026
36
Grok 4.3
62.2569.9318.54$0.061

How this comparison works

Every model tracked across our benchmarks is automatically classified as open-weight (its weights are published and downloadable — for example on Hugging Face) or closed (available only through a provider's own API, with no downloadable weights). This is computed automatically from provider and dataset metadata — not a hand-picked list — so it stays current as new models ship. Note that "open-weight" is not the same as "fully open-source": some open-weight licenses restrict commercial use or large-scale deployment.

Reading the chart

Each dot is one model, positioned by its release date (x-axis) and benchmark score (y-axis). The two stepped lines are the frontier for each group — the running-best score achieved so far by any open-weight model, and separately by any closed model. Because a frontier can only go up, it forms a staircase: it jumps the moment a new best model ships and holds flat until the next one beats it.

What "months behind" means

The headline number is the gap between the two frontiers, measured in time rather than raw score: we take the open-weight frontier's current best score, then find the earliest date the closed frontier reached that same score. The difference between those two dates, in months, is how far behind open-weight currently trails on that benchmark. This mirrors the methodology used by Epoch AI and Stanford HAI's AI Index for tracking the open-vs-closed gap over time.

Why the lag varies by benchmark

The gap is not uniform — some tasks (broad knowledge benchmarks like MMLU) tend to close faster than others (frontier math and long-horizon agentic tasks), where the largest closed labs' scale and compute advantage shows up most. Switch benchmarks above to see how the lag differs across reasoning, coding, and general knowledge. For head-to-head cost comparisons of specific models, see the compare page; for capability-per-dollar across all benchmarks, see the main benchmarks hub.

Frequently asked questions

What counts as an 'open-weight' model?
A model whose trained weights are published and downloadable (for example, hosted on Hugging Face), as opposed to being accessible only through a provider's hosted API. Open-weight does not always mean open-source in the strict license sense — some open-weight releases carry licenses that restrict commercial use above a usage threshold, or non-commercial-only terms.
How is open-weight status determined — is it a manual list?
No manual list. Status is derived automatically from two independent signals: whether a provider's API metadata links the model to a public Hugging Face repository, and Epoch AI's own model-accessibility classification. Both refresh automatically on our daily data pipeline, so new open-weight releases are tagged the day they appear rather than waiting on a hand-curated update.
Why do some models not appear on the chart?
The chart only plots models with a known release date. Some benchmark entries — particularly very old or unusually-named model snapshots — don't have a confidently matched release date in our data sources and are excluded from this specific chart, though they still appear in the leaderboard table below it.
Is the open-weight frontier always behind the closed frontier?
Not necessarily on every benchmark at every moment, but historically yes on most — closed frontier labs have generally reached a given score level first. The gap has been narrowing across most benchmarks over the past few years, which is exactly what this page is built to show over time.