Agent Cost Calculator

What does an agent loop actually cost?

Multi-step agents resend their full history on every step, so cost grows quadratically with step count, not linearly — a single-turn chat calculator will badly underestimate this. Model the real shape instead.

Start from
Cheapest for this loop
Qwen3.7 Flash
$29.99 /mo
#Model$/taskRelative$/month
1
QwenQwen3.7 Flash
$0.0060
$29.99
2
DeepSeekDeepSeek V4 Flash 0731
$0.0094
$46.89
3
GLM 4.7 Flash
$0.0127
$63.67
4
QwenQwen3.5-Flash
$0.0129
$64.40
5
GoogleGemma 4 26B A4B
$0.0142
$70.94
6
GLM 5.3 Flash
$0.0146
$72.98
7
DeepSeekDeepSeek V4 Flash 0423
$0.0146
$73.15
8
GoogleGemma 4 31B
$0.0177
$88.64
9
QwenQwen3.5-9B
$0.0185
$92.46
10
OpenAIGPT-5.6 Luna (batch)
$0.0209
$104
11
OpenAIGPT-5.6 Luna Pro (batch)
$0.0209
$104
12
OpenAIGPT-5.4 Nano (batch)
$0.0210
$105
13
QwenQwen3.6 35B A3B
$0.0225
$112
14
QwenQwen3 Coder Next
$0.0255
$127
15
GoogleGemini 3.1 Flash Lite (batch)
$0.0261
$130
16
QwenQwen3.8 Flash
$0.0290
$145
17
MistralMistral Small 4
$0.0297
$149
18
GoogleGemini 3.5 Flash Lite (batch)
$0.0332
$166
19
GoogleGemini 3.7 Flash (batch)
$0.0382
$191
20
QwenQwen3.6 Flash
$0.0392
$196
21
OpenAIGPT-5.6 Luna
$0.0417
$209
22
OpenAIGPT-5.6 Luna Pro
$0.0417
$209
23
OpenAIGPT-5.4 Nano
$0.0420
$210
24
DeepSeekDeepSeek V4 Flash Vision Exp
$0.0424
$212
25
QwenQwen3.5-27B
$0.0428
$214
26
QwenQwen3.5-35B-A3B
$0.0509
$254
27
GoogleGemini 3.1 Flash Lite
$0.0522
$261
28
GoogleGemini 3.1 Flash Lite Preview
$0.0522
$261
29
QwenQwen3.5 Plus 2026-02-15
$0.0543
$271
30
QwenQwen3.5-122B-A10B
$0.0570
$285
31
QwenQwen3.5 Plus 2026-04-20
$0.0626
$313
32
QwenQwen3.7 Plus
$0.0634
$317
33
GoogleGemini 3.5 Flash Lite
$0.0663
$332
34
QwenQwen3.6 Plus
$0.0678
$339
35
GoogleGemini 3.6 Flash (batch)
$0.0763
$381
36
GoogleGemini 3.7 Flash
$0.0763
$381
37
OpenAIGPT-5.4 Mini (batch)
$0.0783
$391
38
QwenQwen3.5 397B A17B
$0.0814
$407
39
QwenQwen3.8 27B
$0.0887
$444
40
GLM 5
$0.1163
$582
41
MoonshotKimi K2.5
$0.1221
$610
42
QwenQwen3.6 27B
$0.1252
$626
43
MoonshotKimi K2.7 Code
$0.1348
$674
44
GoogleGemini 3.6 Flash
$0.1526
$763
45
GoogleGemini 3.5 Flash (batch)
$0.1565
$783
46
OpenAIGPT-5.4 Mini
$0.1565
$783
47
QwenQwen3 Max Thinking
$0.1587
$793
48
DeepSeekDeepSeek V4 Pro 0423
$0.1632
$816
49
Grok Build 0.1
$0.1876
$938
50
MoonshotKimi K2.6
$0.1893
$946
51
MoonshotKimi K2.7 Code (batch)
$0.1893
$946
52
AnthropicClaude Sonnet 5 (batch)
$0.2034
$1,017
53
OpenAIGPT-5.6 Sol (batch)
$0.2034
$1,017
54
OpenAIGPT-5.6 Sol Pro (batch)
$0.2034
$1,017
55
GoogleGemini 3.1 Pro Preview (batch)
$0.2087
$1,044
56
OpenAIGPT-5.6 Terra (batch)
$0.2087
$1,044
57
OpenAIGPT-5.6 Terra Pro (batch)
$0.2087
$1,044
58
QwenQwen3.6 Max Preview
$0.2144
$1,072
59
DeepSeekDeepSeek V4 Pro 0813
$0.2164
$1,082
60
GLM 5.2
$0.2304
$1,152
61
GLM 5 Turbo
$0.2335
$1,168
62
GLM 5V Turbo
$0.2335
$1,168
63
Grok 4.20
$0.2345
$1,172
64
Grok 4.20 Multi-Agent
$0.2345
$1,172
65
Grok 4.3
$0.2345
$1,172
66
GLM 5.1
$0.2440
$1,220
67
OpenAIGPT-5.4 (batch)
$0.2609
$1,304
68
GLM 5.3
$0.2711
$1,355
69
QwenQwen3.7 Max
$0.2845
$1,422
70
AnthropicClaude Sonnet 4.6 (batch)
$0.3051
$1,526
71
MistralMistral Medium 3.5
$0.3051
$1,526
72
GoogleGemini 3.5 Flash
$0.3131
$1,565
73
OpenAIGPT-5.2-Codex
$0.3838
$1,919
74
OpenAIGPT-5.3-Codex
$0.3838
$1,919
75
QwenQwen3.8 2.4T A95B
$0.3857
$1,929
76
QwenQwen3.8 Max
$0.3857
$1,929
77
Grok 4.5
$0.3857
$1,929
78
Grok 4.6
$0.3857
$1,929
79
AnthropicClaude Sonnet 5
$0.4069
$2,034
80
OpenAIGPT-5.6 Sol
$0.4069
$2,034
81
OpenAIGPT-5.6 Sol Pro
$0.4069
$2,034
82
GoogleGemini 3.1 Pro Preview
$0.4174
$2,087
83
GoogleGemini 3.1 Pro Preview Custom Tools
$0.4174
$2,087
84
OpenAIGPT-5.6 Terra
$0.4174
$2,087
85
OpenAIGPT-5.6 Terra Pro
$0.4174
$2,087
86
AnthropicClaude Opus 4.6 (batch)
$0.5086
$2,543
87
AnthropicClaude Opus 4.7 (batch)
$0.5086
$2,543
88
AnthropicClaude Opus 4.8 (batch)
$0.5086
$2,543
89
AnthropicClaude Opus 5 (batch)
$0.5086
$2,543
90
OpenAIGPT-5.4
$0.5218
$2,609
91
OpenAIGPT-5.5 (batch)
$0.5218
$2,609
92
AnthropicClaude Sonnet 4.6
$0.6103
$3,051
93
MoonshotKimi K3
$0.6103
$3,051
94
AnthropicClaude Fable 5 (batch)
$1.02
$5,086
95
AnthropicClaude Opus 4.6
$1.02
$5,086
96
AnthropicClaude Opus 4.7
$1.02
$5,086
97
AnthropicClaude Opus 4.8
$1.02
$5,086
98
AnthropicClaude Opus 5
$1.02
$5,086
99
OpenAIGPT Chat Latest
$1.04
$5,218
100
OpenAIGPT-5.5
$1.04
$5,218
101
AnthropicClaude Fable 5
$2.03
$10,172
102
AnthropicClaude Opus 4.8 (Fast)
$2.03
$10,172
103
AnthropicClaude Opus 5 (Fast)
$2.03
$10,172
104
OpenAIGPT-5.4 Pro (batch)
$3.13
$15,654
105
OpenAIGPT-5.5 Pro (batch)
$3.13
$15,654
106
AnthropicClaude Opus 4.7 (Fast)
$6.10
$30,515
107
OpenAIGPT-5.4 Pro
$6.26
$31,308
108
OpenAIGPT-5.5 Pro
$6.26
$31,308

Assumes fixed context (system prompt + tool schema) is resent every step, plus accumulating history from prior steps — the real shape of a stateless chat-completions API loop. Retry overhead is applied as a flat multiplier on total cost.

Why agent costs grow quadratically, not linearly

Chat-completion APIs are stateless — every call resends the full conversation, not just the new part. In a tool-using agent loop, each step's response (and the tool's observation) gets appended to that history, so by step N you're resending everything from steps 1 through N-1 on top of the new content. If each step adds roughly T tokens of new context, total input tokens across N steps work out to approximately T × N(N+1)/2 — quadratic, not linear, growth. Double the step count and the token bill doesn't double, it roughly quadruples. This is the actual mechanism behind the commonly-cited "agents use 4-15x more tokens than chat" figures.

What the inputs mean

Fixed context is what's resent unchanged on every single step — your system prompt plus every tool's schema definition. Tool schemas aren't free: real-world measurements put the average around 700 tokens per tool definition, so an agent with a dozen tools can spend a meaningful chunk of its context budget before any real work happens. Context added per step is the growing part — the model's own response plus whatever the tool returned, both of which get resent on every subsequent step. Output tokens per step is only what the model actually generates each turn (billed at the output rate), which is typically smaller than the total context growth per step, since tool observations aren't LLM-generated but still count as input on the next call.

Why the presets differ so much

The four starting presets are grounded in real step-count data rather than round numbers: coding agents commonly run 12-30 steps per task (SWE-bench agent trajectories average in that range), while a simple lookup-and-answer loop might resolve in 3-4. Because cost grows with the square of step count, that difference matters far more than it looks — a coding agent isn't just "4x more steps" expensive than a simple loop, it's closer to 15-20x on the input-token side alone.

What this doesn't model

This is a planning estimate, not a trace replay: real agents don't take a fixed number of steps every time, context isn't always a clean linear accumulation (some frameworks prune or summarize older history), and retry behavior varies a lot by task. Use the retry-overhead slider to sanity-check a range rather than treating the output as exact. For a single-turn (non-agentic) estimate, use the regular calculator instead.

Frequently asked questions

Why does cost grow quadratically instead of linearly with step count?
Because chat-completion APIs are stateless and resend the full conversation on every call. By step N, you're resending everything accumulated from steps 1 through N-1, so total input tokens scale roughly with the square of the step count, not the step count itself.
How much context do tool definitions actually cost?
Real measurements of MCP servers found tool-schema overhead averaging around 700 tokens per tool, with some heavier tools running 1,000+ tokens each — in tool-heavy setups this can consume 30-50% of the context budget before any real work happens.
Is this the same as the regular calculator?
No. The regular calculator models single-turn chat (one request, one response). This models a multi-step agent loop where each step resends the accumulating history, which produces a very different — and much larger — cost shape for the same nominal traffic volume.
Why include a retry/failure overhead slider?
Agent tasks fail and retry more often than single-turn chat calls — tool-call errors, rate limits, and multi-step reasoning failures all add real, billed tokens that a naive per-task estimate misses. The slider applies that as a flat percentage on top of the modeled cost so you can see a realistic range.