Cache Savings

Which model rewards caching the most?

Every provider discounts cached input tokens differently — some by 90%+, some far less. Enter the part of your prompt that repeats across requests and see exactly what caching is worth, model by model.

Start from
Deepest cache discount
DeepSeek V4 Flash Vision Exp97% off cached tokens
$23.00 saved/mo
#ModelDiscountWithout cacheWith cacheSaved / mo
1
DeepSeekDeepSeek V4 Flash Vision Exp
97%$39.60$16.60$23.00
2
DeepSeekDeepSeek V4 Pro 0813
97%$202$84.78$117
3
DeepSeekDeepSeek V4 Pro 0423
92%$157$70.42$86.18
4
GoogleGemini 3.1 Flash Lite (batch)
90%$22.50$10.30$12.20
5
GoogleGemini 3.6 Flash (batch)
90%$67.50$31.00$36.50
6
GoogleGemini 3.7 Flash
90%$67.50$31.00$36.50
7
OpenAIGPT-5.4 Mini (batch)
90%$67.50$31.00$36.50
8
AnthropicClaude Sonnet 5 (batch)
90%$180$82.80$97.20
9
AnthropicClaude Sonnet 4.6 (batch)
90%$270$124$146
10
AnthropicClaude Sonnet 5
90%$360$166$194
11
AnthropicClaude Opus 4.6 (batch)
90%$450$207$243
12
AnthropicClaude Opus 4.7 (batch)
90%$450$207$243
13
AnthropicClaude Opus 4.8 (batch)
90%$450$207$243
14
AnthropicClaude Opus 5 (batch)
90%$450$207$243
15
AnthropicClaude Sonnet 4.6
90%$540$248$292
16
AnthropicClaude Fable 5 (batch)
90%$900$414$486
17
AnthropicClaude Opus 4.6
90%$900$414$486
18
AnthropicClaude Opus 4.7
90%$900$414$486
19
AnthropicClaude Opus 4.8
90%$900$414$486
20
AnthropicClaude Opus 5
90%$900$414$486
21
AnthropicClaude Fable 5
90%$1,800$828$972
22
AnthropicClaude Opus 4.8 (Fast)
90%$1,800$828$972
23
AnthropicClaude Opus 5 (Fast)
90%$1,800$828$972
24
AnthropicClaude Opus 4.7 (Fast)
90%$5,400$2,484$2,916
25
GoogleGemma 4 26B A4B
90%$12.60$5.80$6.80
26
GoogleGemini 3.5 Flash Lite (batch)
90%$27.00$12.42$14.58
27
GoogleGemini 3.1 Flash Lite
90%$45.00$20.70$24.30
28
GoogleGemini 3.1 Flash Lite Preview
90%$45.00$20.70$24.30
29
GoogleGemini 3.5 Flash Lite
90%$54.00$24.84$29.16
30
GoogleGemini 3.5 Flash (batch)
90%$135$62.10$72.90
31
GoogleGemini 3.6 Flash
90%$135$62.10$72.90
32
GoogleGemini 3.1 Pro Preview (batch)
90%$180$82.80$97.20
33
GoogleGemini 3.5 Flash
90%$270$124$146
34
GoogleGemini 3.1 Pro Preview
90%$360$166$194
35
GoogleGemini 3.1 Pro Preview Custom Tools
90%$360$166$194
36
MistralMistral Small 4
90%$27.00$12.42$14.58
37
MistralMistral Medium 3.5
90%$270$124$146
38
MoonshotKimi K3
90%$540$248$292
39
OpenAIGPT-5.4 Nano (batch)
90%$18.00$8.28$9.72
40
OpenAIGPT-5.6 Luna (batch)
90%$18.00$8.28$9.72
41
OpenAIGPT-5.6 Luna Pro (batch)
90%$18.00$8.28$9.72
42
OpenAIGPT-5.4 Nano
90%$36.00$16.56$19.44
43
OpenAIGPT-5.6 Luna
90%$36.00$16.56$19.44
44
OpenAIGPT-5.6 Luna Pro
90%$36.00$16.56$19.44
45
OpenAIGPT-5.4 Mini
90%$135$62.10$72.90
46
OpenAIGPT-5.6 Sol (batch)
90%$180$82.80$97.20
47
OpenAIGPT-5.6 Sol Pro (batch)
90%$180$82.80$97.20
48
OpenAIGPT-5.6 Terra (batch)
90%$180$82.80$97.20
49
OpenAIGPT-5.6 Terra Pro (batch)
90%$180$82.80$97.20
50
OpenAIGPT-5.4 (batch)
90%$225$103$122
51
OpenAIGPT-5.2-Codex
90%$315$145$170
52
OpenAIGPT-5.3-Codex
90%$315$145$170
53
OpenAIGPT-5.6 Sol
90%$360$166$194
54
OpenAIGPT-5.6 Sol Pro
90%$360$166$194
55
OpenAIGPT-5.6 Terra
90%$360$166$194
56
OpenAIGPT-5.6 Terra Pro
90%$360$166$194
57
OpenAIGPT-5.4
90%$450$207$243
58
OpenAIGPT-5.5 (batch)
90%$450$207$243
59
OpenAIGPT Chat Latest
90%$900$414$486
60
OpenAIGPT-5.5
90%$900$414$486
61
OpenAIGPT-5.4 Pro (batch)
90%$2,700$1,242$1,458
62
OpenAIGPT-5.5 Pro (batch)
90%$2,700$1,242$1,458
63
OpenAIGPT-5.4 Pro
90%$5,400$2,484$2,916
64
OpenAIGPT-5.5 Pro
90%$5,400$2,484$2,916
65
QwenQwen3.5-9B
90%$18.00$8.28$9.72
66
QwenQwen3.5 Plus 2026-02-15
90%$46.80$21.53$25.27
67
QwenQwen3.5-122B-A10B
90%$46.80$21.53$25.27
68
QwenQwen3.5 Plus 2026-04-20
90%$54.00$24.84$29.16
69
QwenQwen3.5 397B A17B
90%$70.20$32.29$37.91
70
QwenQwen3 Max Thinking
90%$140$64.58$75.82
71
QwenQwen3.6 Max Preview
90%$185$85.07$99.79
72
GoogleGemini 3.7 Flash (batch)
90%$33.84$15.59$18.25
73
QwenQwen3.6 Flash
90%$33.84$15.59$18.25
74
QwenQwen3.6 Plus
90%$58.50$26.96$31.54
75
QwenQwen3.5-27B
90%$35.10$16.20$18.90
76
QwenQwen3.8 Flash
89%$27.00$12.53$14.47
77
QwenQwen3.5-Flash
89%$11.70$5.44$6.26
78
QwenQwen3.8 2.4T A95B
88%$360$171$189
79
QwenQwen3.8 Max
88%$360$171$189
80
Grok 4.5
85%$360$176$184
81
Grok 4.20
84%$225$112$113
82
Grok 4.20 Multi-Agent
84%$225$112$113
83
Grok 4.3
84%$225$112$113
84
MoonshotKimi K2.5
83%$108$54.00$54.00
85
GLM 4.7 Flash
83%$10.80$5.40$5.40
86
MoonshotKimi K2.6
83%$171$85.68$85.32
87
GLM 5.2
81%$214$110$105
88
GLM 5.1
81%$227$116$111
89
GLM 5.3
81%$252$129$123
90
DeepSeekDeepSeek V4 Flash 0731
80%$9.00$4.68$4.32
91
MoonshotKimi K2.7 Code (batch)
80%$171$88.92$82.08
92
QwenQwen3.7 Flash
80%$5.40$2.81$2.59
93
QwenQwen3.7 Plus
80%$57.60$29.95$27.65
94
QwenQwen3.8 27B
80%$76.50$39.78$36.72
95
QwenQwen3.6 27B
80%$108$56.16$51.84
96
QwenQwen3.7 Max
80%$266$138$127
97
Grok Build 0.1
80%$180$93.60$86.40
98
GLM 5.3 Flash
80%$13.50$7.02$6.48
99
GLM 5
80%$108$56.16$51.84
100
GLM 5 Turbo
80%$216$112$104
101
GLM 5V Turbo
80%$216$112$104
102
DeepSeekDeepSeek V4 Flash 0423
79%$14.04$7.34$6.70
103
Grok 4.6
75%$360$198$162
104
MoonshotKimi K2.7 Code
73%$119$66.96$51.84
105
QwenQwen3.6 35B A3B
50%$18.00$12.60$5.40
106
GoogleGemma 4 31B
44%$16.20$11.88$4.32
107
QwenQwen3 Coder Next
42%$21.60$16.20$5.40
108
QwenQwen3.5-35B-A3B
0%$45.00$45.00$0.00

Figures cover only the repeated-context tokens above — not your full bill. See the calculator to project total monthly spend.

Why the discount varies so much by provider

Prompt caching works the same way everywhere in principle — a provider stores the computed state for a prompt prefix so a later request that starts with the same prefix skips recomputing it — but the discount each provider actually passes on to you is a business decision, not a law of physics, and it varies widely. Some providers price a cache hit near-zero; others give a smaller cut. That gap changes which model is the better choice once your traffic has a large repeated prefix, even if the two models’ standard input prices look similar.

What counts as "repeated context"

Only the part of a prompt that’s byte-for-byte identical across requests benefits — a fixed system prompt, a set of few-shot examples, a tool/function schema, or a retrieved document set that doesn’t change between calls. The variable part (the user’s actual message) never benefits from caching, since it’s different every time. Enter only the repeated portion above; mixing in the variable part will overstate your savings.

How the numbers here are calculated

Discount % is 1 − (cached rate ÷ standard input rate), using each model’s own live rates from our pricing table. Without cache assumes every request pays the full input rate for the repeated tokens; with cache blends the hit rate you set — at a 60% hit rate, 60% of repeated-context tokens bill at the cached rate and 40% at the standard rate, matching how a real cache miss (cold start, expired TTL, or a session outside the cache window) actually bills.

What this doesn't model

Some providers charge more than the standard rate the first time a prefix is cached (a "cache write"), before later reads become cheap — this calculator only models the discounted read side, since write costs are typically a small fraction of total spend once a prefix is reused many times. It also doesn’t include your output tokens or the variable part of each prompt; use the calculator for a full monthly bill projection.

Frequently asked questions

Which providers give the deepest prompt-caching discount?
It varies by provider and changes over time as pricing updates — that's exactly why this page computes it live from each model's current cached-input rate rather than stating a fixed number. Check the ranked table above for the current figures.
Why is my real bill different from this estimate?
This tool only prices the repeated-context tokens you enter, not your full request (it excludes the variable part of each prompt and all output tokens), and it doesn't model any cache-write premium some providers charge on the first use of a prefix. Use the calculator for a complete monthly projection.
What should I set the cache hit rate to?
It depends on your traffic pattern and the provider's cache time-to-live. Bursty traffic that reuses a prefix within a short window (minutes) typically sees a high hit rate; sparse or bursty-then-idle traffic sees a lower one. If you're unsure, run the numbers at a conservative 40-50% and a high 80-90% to see the range.
Are the values I enter sent anywhere?
No. All calculation happens in your browser. The numbers you type are never transmitted to or stored on our servers.