← All guides

How LLM API Prices Actually Changed: A Data-Driven Look at Our Own Tracking Log

By TokenCost Editorial · Published Aug 2026

Most "LLM prices are falling" claims are vibes, not data — a handful of headline launches cited as if they represent the whole market. We don't have to guess: our price-trends log has been diffing real, dated pricing snapshots since July 13, 2026, which means we can actually answer the question with numbers instead of anecdotes. This is what those numbers say, updated as the log grows.

The topline: cuts outnumber increases, but not by as much as you'd think

Over the six-week tracking window behind this article, our pipeline logged 559 distinct events — price changes, new model launches, and retirements. Filtering to genuine headline pricing moves (input and output rates, excluding cached-input rates for reasons explained below):

  • 177 price cuts vs. 142 price increases across input and output rates combined
  • The median cut was -22% on input, -27% on output
  • The median increase was smaller in most cases but still substantial — +47% on input, +45% on output — meaning individual increases, when they happen, aren't small corrections

The honest framing: prices trend down more often than up, but any given model, in any given week, has a real chance of getting more expensive — repricing runs in both directions constantly, not just downward at launch and then flat forever.

The biggest single moves

Steepest input-price cuts:

ModelProviderChange
GPT-5.6 Luna ProOpenAI$0.50 → $0.10 (-80%)
GPT-5.6 LunaOpenAI$0.50 → $0.10 (-80%)
DeepSeek V4 Pro 0423DeepSeek$1.60 → $0.465 (-70.9%)

Steepest input-price increases:

ModelProviderChange
GLM 5.2Z.AI$0.098 → $0.76 (+675.5%)
GLM 5.2Z.AI$0.284 → $1.19 (+319.0%)
GLM 5.2Z.AI$0.308 → $1.19 (+286.4%)

GLM 5.2 shows up repeatedly on the increases side alone — cut once, then re-priced upward three separate times within the tracking window. That's not a typo; it's a real pattern worth naming: some providers reprice repeatedly and sharply in short windows, especially newer entrants still finding a stable price point, rather than settling on a rate and holding it for a quarter or a year the way the largest labs tend to. DeepSeek V4 Pro shows the same volatility from the other direction — a steep cut in late August followed by a separate, even steeper output-price increase days later.

By provider: who cuts, who raises

ProviderCutsIncreases
Z.AI5028
OpenAI4812
DeepSeek4856
Qwen4757
Moonshot4436
Google2726
xAI10

OpenAI is the most cut-skewed of the high-volume providers — 4 cuts for every 1 increase in this window. Qwen and DeepSeek are the mirror image: both show more increases than cuts, the only two providers in our tracking where that's true. Neither pattern is guaranteed to hold going forward — this reflects one real, measured window, not a permanent law — but it's actual measured behavior, not a guess.

The cache-rate caveat

We deliberately excluded cached-input pricing from the headline numbers above, and it's worth explaining why: cached rates are usually small numbers (often under $0.10/1M), so a modest absolute change produces an enormous percentage swing that can dominate a naive "biggest price changes" list without reflecting the market at all. Some of that volatility is also a data-quality artifact rather than a market move: when a provider publishes a real cached-input rate for the first time, our pipeline switches from an estimated fallback to the real number (see methodology), which shows up in the log as a "price change" even though nothing in the market actually moved. We track it, but we don't headline it, for exactly that reason.

New models vs. retirements

Sixty-three new models entered our tracking during this window against sixteen retirements — roughly a 4:1 ratio of launches to sunsets, concentrated among OpenAI (16 new), Anthropic (15 new), and Google (14 new). That pace is itself a cost-planning signal: if you're pricing a product around "the current cheapest model in tier X," expect that answer to have a shelf life measured in weeks, not quarters. The cheapest models page and price trends log reflect this daily rather than requiring you to re-research it.

Why this data is different from most "LLM pricing trend" content

Nearly everything else written about LLM price trends is a snapshot compared to a memory — "GPT-4 cost $X in 2023, now it's $Y" — which is true but tells you nothing about the path between those two points, or how often the price actually moved along the way. Our log is a genuine diff, reconstructed from real git-committed pricing snapshots and extended by a daily automated pipeline since — every event above is a real, dated change we captured, not an estimate or a reconstruction from press releases. See the full, filterable live log for the complete and continuously updating picture; the figures on this page are a snapshot as of the date above and will drift from the live log as new events accumulate.