Head to head

Put the contenders side by side.

Add up to four columns and pick any provider's model in each. Line up pricing, context window, and a shared cost example — the best value in each row is called out automatically.

Comparing 3 of 4 columns
Model
Standard
Provider
Model
Budget
Provider
Model
Budget
Provider
Model
Input priceper 1M tokens
$1.00
$0.14
$0.07BEST
Output priceper 1M tokens
$5.00
$0.28BEST
$0.34
Cached inputper-model rate
$0.10
$0.03
$0.01
Context windowmax tokens
1MBEST
1MBEST
262K
Cost example1M in + 250K out
$2.25
$0.21
$0.16BEST

Cost example assumes 1M input + 250K output tokens at each model's standard (uncached) rate. Model your real traffic on the calculator.

How to choose between models

A cheaper headline price isn’t always the cheaper model for your workload. Which rate matters most depends on the shape of your traffic, so the comparison lines up the numbers that actually drive a bill:

  • Input vs output price. Which rate matters more depends on your traffic shape — see Input vs Output Tokens for why output usually costs more and how to weigh the two.
  • Cached input. If you reuse a big fixed prompt or document across many calls, a model with an aggressive cached-input discount can beat a nominally cheaper model that lacks one — see Prompt Caching Explained.
  • Context window. A larger window lets you pass more context in one call, but remember you pay for every token you send — a big window is a capability, not a discount.
  • Cost example. The final row prices a fixed, realistic mix — 1M input + 250K output tokens at each model’s standard (uncached) rate — so you get one comparable dollar figure per model. The best value in every row is highlighted automatically.

Add up to four columns, pick any provider’s model in each, and read down the rows. When you’ve shortlisted, model your real traffic on the calculator, or weigh capability against price on the benchmarks. Full rates for every model live on the pricing page.

Frequently asked questions

Which model is cheapest overall?
There is no single answer — it depends on your input/output ratio and cache usage. A model with cheap input but expensive output wins for short-answer workloads and loses for long-generation ones. Use the cost-example row for a like-for-like figure, then confirm with the calculator using your real traffic.
What does the 'cost example' row assume?
A fixed 1M input tokens plus 250K output tokens, priced at each model’s standard (uncached) input and output rates. It is a shared yardstick for comparison, not a prediction of your bill.
How many models can I compare at once?
Up to four columns on desktop and three on mobile, each independently selectable by provider and model. The best value in each attribute row is flagged automatically.
Does a bigger context window cost more?
The window itself is a capacity limit, not a charge. But you pay per token for everything you actually send, so filling a large window with context increases the input cost of each call.