Open-weight vs closed models: the race, charted.
Open-weight releases — Llama, DeepSeek, Qwen, GLM, Kimi, Gemma and others — are closing in on closed, API-only frontier models. This page plots both groups by release date and benchmark score, and computes exactly how many months behind the closed frontier the best open-weight model currently sits, per benchmark.
LiveBench — open-weight vs closed, over time
Each dot is a model release, plotted by its release date and score. The stepped lines trace the running-best ("frontier") for each group — 8 open-weight, 28 closed models tracked.
How this comparison works
Every model tracked across our benchmarks is automatically classified as open-weight (its weights are published and downloadable — for example on Hugging Face) or closed (available only through a provider's own API, with no downloadable weights). This is computed automatically from provider and dataset metadata — not a hand-picked list — so it stays current as new models ship. Note that "open-weight" is not the same as "fully open-source": some open-weight licenses restrict commercial use or large-scale deployment.
Reading the chart
Each dot is one model, positioned by its release date (x-axis) and benchmark score (y-axis). The two stepped lines are the frontier for each group — the running-best score achieved so far by any open-weight model, and separately by any closed model. Because a frontier can only go up, it forms a staircase: it jumps the moment a new best model ships and holds flat until the next one beats it.
What "months behind" means
The headline number is the gap between the two frontiers, measured in time rather than raw score: we take the open-weight frontier's current best score, then find the earliest date the closed frontier reached that same score. The difference between those two dates, in months, is how far behind open-weight currently trails on that benchmark. This mirrors the methodology used by Epoch AI and Stanford HAI's AI Index for tracking the open-vs-closed gap over time.
Why the lag varies by benchmark
The gap is not uniform — some tasks (broad knowledge benchmarks like MMLU) tend to close faster than others (frontier math and long-horizon agentic tasks), where the largest closed labs' scale and compute advantage shows up most. Switch benchmarks above to see how the lag differs across reasoning, coding, and general knowledge. For head-to-head cost comparisons of specific models, see the compare page; for capability-per-dollar across all benchmarks, see the main benchmarks hub.
