Put the contenders side by side.
Add up to four columns and pick any provider's model in each. Line up pricing, context window, and a shared cost example — the best value in each row is called out automatically.
Cost example assumes 1M input + 250K output tokens at each model's standard (uncached) rate. Model your real traffic on the calculator.
How to choose between models
A cheaper headline price isn’t always the cheaper model for your workload. Which rate matters most depends on the shape of your traffic, so the comparison lines up the numbers that actually drive a bill:
- Input vs output price. Which rate matters more depends on your traffic shape — see Input vs Output Tokens for why output usually costs more and how to weigh the two.
- Cached input. If you reuse a big fixed prompt or document across many calls, a model with an aggressive cached-input discount can beat a nominally cheaper model that lacks one — see Prompt Caching Explained.
- Context window. A larger window lets you pass more context in one call, but remember you pay for every token you send — a big window is a capability, not a discount.
- Cost example. The final row prices a fixed, realistic mix — 1M input + 250K output tokens at each model’s standard (uncached) rate — so you get one comparable dollar figure per model. The best value in every row is highlighted automatically.
Add up to four columns, pick any provider’s model in each, and read down the rows. When you’ve shortlisted, model your real traffic on the calculator, or weigh capability against price on the benchmarks. Full rates for every model live on the pricing page.
