The cheapest LLM APIs, ranked for your workload.
“Cheapest” depends on whether you send more than you generate. Pick your workload mix and every model re-ranks by the blended price you’ll actually pay.
50% input · 50% output. Ranked by blended price per 1M tokens for that mix.
Blended price = input rate × input share + output rate × output share, per 1M tokens. Refreshed daily. Full input/output/cached rates are on the pricing page.
How to find the cheapest model for your use case
The lowest headline price isn’t always the cheapest model to run. Because input and output tokens are priced separately — and output usually costs several times more — the winner flips depending on your workload:
- Input-heavy (80/20) — retrieval-augmented generation, classification, extraction, routing. You send a lot and generate a little, so the input rate dominates.
- Output-heavy (20/80) — drafting, long-form writing, code generation. The output rate dominates, so a low output price matters most.
- Balanced (50/50) — chat and general assistants, a reasonable default when you’re unsure.
The blended figure reweights each model’s input and output rates for the mix you pick, so the ranking reflects what you’ll really pay. Once you’ve shortlisted, sanity-check capability against price on the benchmarks — the very cheapest model isn’t a bargain if it can’t do the job — and project a full monthly bill on the calculator. To price a specific prompt, use the token counter. For the full breakdown of why the ranking depends on your ratio, see Which Frontier Model Is Actually Cheapest for Your Workload? — and for ways to cut cost without switching models at all, see How to Cut Your LLM API Bill by 80%.
