Bill Projector

What will this cost me to run?

Enter your expected traffic and see the monthly bill across every model — before a single user signs up.

Start from
BudgetQwen
Qwen3.7 Flash
$0.03 in · $0.13 out / 1M
$119
per month · $0.02 / user
StandardxAI
Grok Build 0.1
$1 in · $2 out / 1M
$2,568
per month · $0.51 / user
PremiumAnthropic
Claude Fable 5 (batch)
$5 in · $25 out / 1M
$21,570
per month · $4.31 / user
#ModelIn / OutRelativeMonthly
1
Qwen3.7 FlashCHEAPEST
$0.03/$0.13
$119
2
Qwen3.5-9B
$0.1/$0.15
$221
3
Qwen3.5-Flash
$0.065/$0.26
$242
4
Gemma 4 26B A4B
$0.07/$0.34
$296
5
GLM 4.7 Flash
$0.06/$0.4
$321
6
DeepSeek V4 Flash
$0.14/$0.28
$360
7
Gemma 4 31B
$0.14/$0.4
$424
8
GPT-5.4 Nano (batch)
$0.1/$0.625
$506
9
Mistral Small 4
$0.15/$0.6
$557
10
Gemini 3.1 Flash Lite (batch)
$0.125/$0.75
$614
11
Qwen3.5-35B-A3B
$0.14/$1
$784
12
Qwen3.6 35B A3B
$0.14/$1
$784
13
Qwen3 Coder Next
$0.18/$0.9
$786
14
Qwen3.6 Flash
$0.188/$1.125
$922
15
Gemini 3.5 Flash Lite (batch)
$0.15/$1.25
$947
16
GPT-5.4 Nano
$0.2/$1.25
$1,013
17
DeepSeek V4 Pro
$0.435/$0.87
$1,072
18
Qwen3.5-27B
$0.195/$1.56
$1,193
19
Qwen3.7 Plus
$0.32/$1.28
$1,206
20
Gemini 3.1 Flash Lite
$0.25/$1.5
$1,229
21
Gemini 3.1 Flash Lite Preview
$0.25/$1.5
$1,229
22
Qwen3.5 Plus 2026-02-15
$0.26/$1.56
$1,278
23
Qwen3.5 Plus 2026-04-20
$0.3/$1.8
$1,474
24
Qwen3.5-122B-A10B
$0.26/$2.08
$1,590
25
Qwen3.6 Plus
$0.325/$1.95
$1,597
26
Qwen3.6 27B
$0.3/$2
$1,659
27
GPT-5.4 Mini (batch)
$0.375/$2.25
$1,842
28
Gemini 3.5 Flash Lite
$0.3/$2.5
$1,894
29
Qwen3.5 397B A17B
$0.39/$2.34
$1,916
30
GLM 5.2
$0.749/$2.354
$2,431
31
GPT-5.6 Luna
$0.5/$3
$2,457
32
GPT-5.6 Luna Pro
$0.5/$3
$2,457
33
Kimi K2.5
$0.57/$2.85
$2,480
34
Kimi K2.6
$0.646/$2.72
$2,505
35
Grok Build 0.1
$1/$2
$2,568
36
GLM 5
$0.95/$2.55
$2,835
37
Kimi K2.7 Code
$0.73/$3.5
$3,101
38
GLM 5.1
$0.966/$3.036
$3,135
39
Grok 4.20
$1.25/$2.5
$3,183
40
Grok 4.20 Multi-Agent
$1.25/$2.5
$3,183
41
Grok 4.3
$1.25/$2.5
$3,183
42
Gemini 3.6 Flash (batch)
$0.75/$3.75
$3,236
43
Qwen3 Max Thinking
$0.78/$3.9
$3,365
44
Gemini 3.5 Flash (batch)
$0.75/$4.5
$3,686
45
GPT-5.4 Mini
$0.75/$4.5
$3,686
46
GLM 5 Turbo
$1.2/$4
$4,042
47
GLM 5V Turbo
$1.2/$4
$4,042
48
Claude Sonnet 5 (batch)
$1/$5
$4,314
49
Qwen3.7 Max
$1.475/$4.425
$4,673
50
Gemini 3.1 Pro Preview (batch)
$1/$6
$4,914
51
Qwen3.6 Max Preview
$1.027/$6.162
$5,047
52
GPT-5.4 (batch)
$1.25/$7.5
$6,143
53
GPT-5.6 Terra
$1.25/$7.5
$6,143
54
GPT-5.6 Terra Pro
$1.25/$7.5
$6,143
55
Grok 4.5
$2/$6
$6,282
56
Gemini 3.6 Flash
$1.5/$7.5
$6,471
57
Mistral Medium 3.5
$1.5/$7.5
$6,471
58
Gemini 3.5 Flash
$1.5/$9
$7,371
59
Claude Sonnet 5
$2/$10
$8,628
60
Gemini 3.1 Pro Preview
$2/$12
$9,828
61
Gemini 3.1 Pro Preview Custom Tools
$2/$12
$9,828
62
GPT-5.2-Codex
$1.75/$14
$10,700
63
GPT-5.3 Chat
$1.75/$14
$10,700
64
GPT-5.3-Codex
$1.75/$14
$10,700
65
Claude Opus 4.6 (batch)
$2.5/$12.5
$10,785
66
Claude Opus 4.7 (batch)
$2.5/$12.5
$10,785
67
Claude Opus 4.8 (batch)
$2.5/$12.5
$10,785
68
GPT-5.4
$2.5/$15
$12,285
69
GPT-5.5 (batch)
$2.5/$15
$12,285
70
Claude Sonnet 4.6
$3/$15
$12,942
71
Kimi K3
$3/$15
$12,942
72
Claude Fable 5 (batch)
$5/$25
$21,570
73
Claude Opus 4.6
$5/$25
$21,570
74
Claude Opus 4.7
$5/$25
$21,570
75
Claude Opus 4.8
$5/$25
$21,570
76
Claude Opus 5
$5/$25
$21,570
77
GPT Chat Latest
$5/$30
$24,570
78
GPT-5.5
$5/$30
$24,570
79
GPT-5.6 Sol
$5/$30
$24,570
80
GPT-5.6 Sol Pro
$5/$30
$24,570
81
Claude Fable 5
$10/$50
$43,140
82
Claude Opus 4.8 (Fast)
$10/$50
$43,140
83
Claude Opus 5 (Fast)
$10/$50
$43,140
84
Claude Opus 4.7 (Fast)
$30/$150
$129,420
85
GPT-5.4 Pro
$30/$180
$147,420
86
GPT-5.5 Pro
$30/$180
$147,420

Assumes 30 days/month. Cached input billed at each model's own cached rate. See full rates on the pricing page.

How the calculator works

The projection is a straightforward token-volume model. From your inputs we derive monthly request volume and token counts, then price them against every model’s live rates:

  • Prompts per month = monthly active users × prompts per user per day × 30 days.
  • Input tokens = prompts per month × average input tokens per prompt.
  • Output tokens = prompts per month × average output tokens per prompt.
  • Monthly cost = (input tokens ÷ 1M) × blended input rate + (output tokens ÷ 1M) × output rate, where the blended input rate applies your prompt-cache hit rate: cached tokens are billed at each model’s own cached-input rate and the rest at the standard input rate.

In other words, a 30% cache hit rate means 30% of your input tokens are billed at the model’s cached rate and 70% at the full input rate. Output tokens are never cached, so they’re always billed at the standard output rate.

A worked example

Say you have 5,000 monthly active users, each sending 8 prompts a day, with 1,500 input tokens and 500 output tokens per prompt, at a 30% cache hit rate. That’s 1.2M prompts/month → 1.8B input tokens and 600M output tokens. For a model priced at $2 input / $10 output / $0.20 cached per 1M, the blended input rate is ($2 × 0.7) + ($0.20 × 0.3) = $1.46/1M, so input ≈ 1,800 × $1.46 = $2,628 and output ≈ 600 × $10 = $6,000 — about $8,628/month. Change any slider and the whole table re-ranks instantly.

If this estimate looks higher than you'd like, four concrete levers move it the most — caching, batching, model routing, and context trimming — each with worked math in How to Cut Your LLM API Bill by 80%.

The estimate assumes a 30-day month and steady traffic; real bills vary with burst patterns, retries, and tool-use overhead. Treat it as a planning figure and confirm live rates on the pricing page. See methodology for data sources.

Frequently asked questions

Is the calculator accurate for my real usage?
It is a planning estimate. It models steady monthly traffic at a 30-day month using live per-token rates. Real bills vary with traffic bursts, retries, tool-use overhead, and streaming, so treat the figure as a well-grounded projection rather than an invoice.
How is the prompt-cache hit rate applied?
Only to input tokens. At a 30% hit rate, 30% of input tokens are billed at the model’s cached-input rate and 70% at the standard input rate. Output tokens are never cached and are always billed at the full output rate.
Are the values I enter sent anywhere?
No. All calculation happens in your browser. The traffic and token numbers you type are never transmitted to or stored on our servers.
Why does the cheapest model change when I adjust output length?
Because models have very different input-to-output price ratios. A model with cheap input but expensive output can win for short answers and lose for long ones, so the ranking re-sorts as you change the input/output mix.
What counts as a 'prompt'?
One request/response round trip to the model. Multi-turn conversations send the prior history as input each turn, so long chats increase input tokens — account for that in your average input-tokens figure.