Budget finder

The cheapest LLM APIs, ranked for your workload.

“Cheapest” depends on whether you send more than you generate. Pick your workload mix and every model re-ranks by the blended price you’ll actually pay.

Workload

50% input · 50% output. Ranked by blended price per 1M tokens for that mix.

#ModelProviderInputOutputBlended /1M
1
Qwen3.7 FlashCHEAPESTBudget
Qwen$0.03$0.13$0.08
2
Qwen3.5-9BBudget
Qwen$0.10$0.15$0.13
3
Qwen3.5-FlashBudget
Qwen$0.07$0.26$0.16
4
Gemma 4 26B A4B Budget
Google$0.07$0.34$0.21
5
DeepSeek V4 FlashBudget
DeepSeek$0.14$0.28$0.21
6
GLM 4.7 FlashBudget
Z.AI$0.06$0.40$0.23
7
Gemma 4 31BBudget
Google$0.14$0.40$0.27
8
GPT-5.4 Nano (batch)Budget
OpenAI$0.10$0.63$0.36
9
Mistral Small 4Budget
Mistral$0.15$0.60$0.38
10
Gemini 3.1 Flash Lite (batch)Budget
Google$0.13$0.75$0.44
11
Qwen3 Coder NextBudget
Qwen$0.18$0.90$0.54
12
Qwen3.5-35B-A3BBudget
Qwen$0.14$1.00$0.57
13
Qwen3.6 35B A3BBudget
Qwen$0.14$1.00$0.57
14
DeepSeek V4 ProBudget
DeepSeek$0.43$0.87$0.65
15
Qwen3.6 FlashBudget
Qwen$0.19$1.13$0.66
16
Gemini 3.5 Flash Lite (batch)Budget
Google$0.15$1.25$0.70
17
GPT-5.4 NanoBudget
OpenAI$0.20$1.25$0.72
18
Qwen3.7 PlusBudget
Qwen$0.32$1.28$0.80
19
Gemini 3.1 Flash LiteBudget
Google$0.25$1.50$0.88
20
Gemini 3.1 Flash Lite PreviewBudget
Google$0.25$1.50$0.88
21
Qwen3.5-27BBudget
Qwen$0.20$1.56$0.88
22
Qwen3.5 Plus 2026-02-15Budget
Qwen$0.26$1.56$0.91
23
Qwen3.5 Plus 2026-04-20Budget
Qwen$0.30$1.80$1.05
24
Qwen3.6 PlusBudget
Qwen$0.33$1.95$1.14
25
Qwen3.6 27BBudget
Qwen$0.30$2.00$1.15
26
Qwen3.5-122B-A10BBudget
Qwen$0.26$2.08$1.17
27
GPT-5.4 Mini (batch)Budget
OpenAI$0.38$2.25$1.31
28
Qwen3.5 397B A17BBudget
Qwen$0.39$2.34$1.36
29
Gemini 3.5 Flash LiteBudget
Google$0.30$2.50$1.40
30
Grok Build 0.1Standard
xAI$1.00$2.00$1.50
31
GLM 5.2Budget
Z.AI$0.75$2.35$1.55
32
Kimi K2.6Budget
Moonshot$0.65$2.72$1.68
33
Kimi K2.5Budget
Moonshot$0.57$2.85$1.71
34
GPT-5.6 LunaBudget
OpenAI$0.50$3.00$1.75
35
GPT-5.6 Luna ProBudget
OpenAI$0.50$3.00$1.75
36
GLM 5Budget
Z.AI$0.95$2.55$1.75
37
Grok 4.20Standard
xAI$1.25$2.50$1.88
38
Grok 4.20 Multi-AgentStandard
xAI$1.25$2.50$1.88
39
Grok 4.3Standard
xAI$1.25$2.50$1.88
40
GLM 5.1Budget
Z.AI$0.97$3.04$2.00
41
Kimi K2.7 CodeBudget
Moonshot$0.73$3.50$2.12
42
Gemini 3.6 Flash (batch)Budget
Google$0.75$3.75$2.25
43
Qwen3 Max ThinkingBudget
Qwen$0.78$3.90$2.34
44
GLM 5 TurboStandard
Z.AI$1.20$4.00$2.60
45
GLM 5V TurboStandard
Z.AI$1.20$4.00$2.60
46
Gemini 3.5 Flash (batch)Budget
Google$0.75$4.50$2.63
47
GPT-5.4 MiniBudget
OpenAI$0.75$4.50$2.63
48
Qwen3.7 MaxStandard
Qwen$1.48$4.42$2.95
49
Claude Sonnet 5 (batch)Standard
Anthropic$1.00$5.00$3.00
50
Gemini 3.1 Pro Preview (batch)Standard
Google$1.00$6.00$3.50
51
Qwen3.6 Max PreviewStandard
Qwen$1.03$6.16$3.59
52
Grok 4.5Standard
xAI$2.00$6.00$4.00
53
GPT-5.4 (batch)Standard
OpenAI$1.25$7.50$4.38
54
GPT-5.6 TerraStandard
OpenAI$1.25$7.50$4.38
55
GPT-5.6 Terra ProStandard
OpenAI$1.25$7.50$4.38
56
Gemini 3.6 FlashStandard
Google$1.50$7.50$4.50
57
Mistral Medium 3.5Standard
Mistral$1.50$7.50$4.50
58
Gemini 3.5 FlashStandard
Google$1.50$9.00$5.25
59
Claude Sonnet 5Standard
Anthropic$2.00$10.00$6.00
60
Gemini 3.1 Pro PreviewStandard
Google$2.00$12.00$7.00
61
Gemini 3.1 Pro Preview Custom ToolsStandard
Google$2.00$12.00$7.00
62
Claude Opus 4.6 (batch)Standard
Anthropic$2.50$12.50$7.50
63
Claude Opus 4.7 (batch)Standard
Anthropic$2.50$12.50$7.50
64
Claude Opus 4.8 (batch)Standard
Anthropic$2.50$12.50$7.50
65
GPT-5.2-CodexStandard
OpenAI$1.75$14.00$7.88
66
GPT-5.3 ChatStandard
OpenAI$1.75$14.00$7.88
67
GPT-5.3-CodexStandard
OpenAI$1.75$14.00$7.88
68
GPT-5.4Standard
OpenAI$2.50$15.00$8.75
69
GPT-5.5 (batch)Standard
OpenAI$2.50$15.00$8.75
70
Claude Sonnet 4.6Standard
Anthropic$3.00$15.00$9.00
71
Kimi K3Standard
Moonshot$3.00$15.00$9.00
72
Claude Fable 5 (batch)Premium
Anthropic$5.00$25.00$15.00
73
Claude Opus 4.6Premium
Anthropic$5.00$25.00$15.00
74
Claude Opus 4.7Premium
Anthropic$5.00$25.00$15.00
75
Claude Opus 4.8Premium
Anthropic$5.00$25.00$15.00
76
Claude Opus 5Premium
Anthropic$5.00$25.00$15.00
77
GPT Chat LatestPremium
OpenAI$5.00$30.00$17.50
78
GPT-5.5Premium
OpenAI$5.00$30.00$17.50
79
GPT-5.6 SolPremium
OpenAI$5.00$30.00$17.50
80
GPT-5.6 Sol ProPremium
OpenAI$5.00$30.00$17.50
81
Claude Fable 5Premium
Anthropic$10.00$50.00$30.00
82
Claude Opus 4.8 (Fast)Premium
Anthropic$10.00$50.00$30.00
83
Claude Opus 5 (Fast)Premium
Anthropic$10.00$50.00$30.00
84
Claude Opus 4.7 (Fast)Premium
Anthropic$30.00$150.00$90.00
85
GPT-5.4 ProPremium
OpenAI$30.00$180.00$105.00
86
GPT-5.5 ProPremium
OpenAI$30.00$180.00$105.00

Blended price = input rate × input share + output rate × output share, per 1M tokens. Refreshed daily. Full input/output/cached rates are on the pricing page.

How to find the cheapest model for your use case

The lowest headline price isn’t always the cheapest model to run. Because input and output tokens are priced separately — and output usually costs several times more — the winner flips depending on your workload:

  • Input-heavy (80/20) — retrieval-augmented generation, classification, extraction, routing. You send a lot and generate a little, so the input rate dominates.
  • Output-heavy (20/80) — drafting, long-form writing, code generation. The output rate dominates, so a low output price matters most.
  • Balanced (50/50) — chat and general assistants, a reasonable default when you’re unsure.

The blended figure reweights each model’s input and output rates for the mix you pick, so the ranking reflects what you’ll really pay. Once you’ve shortlisted, sanity-check capability against price on the benchmarks — the very cheapest model isn’t a bargain if it can’t do the job — and project a full monthly bill on the calculator. To price a specific prompt, use the token counter. For the full breakdown of why the ranking depends on your ratio, see Which Frontier Model Is Actually Cheapest for Your Workload? — and for ways to cut cost without switching models at all, see How to Cut Your LLM API Bill by 80%.

Frequently asked questions

What is the cheapest LLM API right now?
It depends on your workload. For output-heavy tasks, the model with the lowest output price usually wins; for input-heavy tasks, the lowest input price wins. Use the workload toggle above to rank models by the blended price for your mix — the top row is the cheapest for that scenario, refreshed daily.
What does 'blended price' mean?
A single per-1M-token figure that combines a model’s input and output rates weighted by your workload mix. For a 50/50 balanced mix it is the average of the input and output rates; for an 80/20 input-heavy mix it leans toward the input rate.
Is the cheapest model always the best choice?
No. A cheap model that can’t complete your task reliably costs more in retries and engineering time. Weigh price against capability on the benchmarks page, and validate on your own tasks before committing.
How often do the rankings update?
Prices refresh automatically every day from live provider data, so the cheapest-first ranking reflects current rates. Always confirm the final price on the provider’s own page before committing spend.