Qwen API pricing
18 Qwen models tracked, updated daily. Full rates below, with guidance on which tier fits which workload.
Cached input rates and full detail on the pricing page. Prices in USD per 1M tokens.
Qwen (Alibaba) runs the largest and most fragmented lineup of any provider we track — over a dozen variants spanning different parameter counts, a dedicated coding model, and a "Thinking" reasoning-focused release, almost entirely in the Budget tier with a couple of models stepping up to Standard. See the live table above for the full current list and rates.
Making sense of the naming
Qwen's model names encode two things at once: a version number (3.5, 3.6, 3.7) and a size/architecture descriptor (parameter count, or an "A_B" mixture-of-experts notation like "35B-A3B", meaning a 35-billion-parameter model with roughly 3 billion active per token). Newer version numbers generally mean improved capability at a similar size class; the size/architecture descriptor is what actually drives the price and speed trade-off within a version.
How to pick from a dozen-plus options without guessing
Given the sheer number of variants, a size-first approach works better than trying to track every release individually:
- Smaller variants (lower parameter counts, or MoE models with fewer active parameters) are the cheapest and fastest — well suited to high-volume classification, tagging, and simple chat.
- Mid-size variants balance cost and general capability — a reasonable default for most chat and moderate coding tasks.
- "Max" and "Plus" branded models step up to Standard-tier pricing and capability, positioned similarly to a Premium option within an otherwise Budget-tier family.
- "Qwen3 Coder" is specifically tuned for code generation — check it against general-purpose variants directly for coding workloads on the SWE-bench and Aider Polyglot benchmarks rather than assuming a larger general model wins by default.
- "Thinking"-branded variants are tuned for extended reasoning before answering — useful for tasks where accuracy on multi-step logic matters more than response latency, at a corresponding cost in output tokens (see Input vs Output Tokens for why reasoning-heavy output changes the cost calculus).
Why Qwen dominates budget rankings
Because so much of the lineup sits in the Budget tier, Qwen models frequently take multiple spots at the top of the cheapest models ranking, especially for output-heavy workloads where a low output rate compounds. If your workload is cost-sensitive and doesn't require frontier-level reasoning, it's worth checking several Qwen size variants against each other before looking elsewhere — the within-family price spread alone can be significant.
What to verify
With this many variants, confirm you're comparing the specific model you intend to deploy against — a size/version mismatch between what you tested and what you ship is an easy mistake in a lineup this large. Re-check the pricing table directly by exact model name before finalizing.
When Qwen vs. a competitor
For pure Budget-tier cost comparisons, check Qwen against DeepSeek and Moonshot directly. For workloads that need the Standard-tier "Max"/"Plus" capability, compare those specifically against Standard-tier models from OpenAI, Google, and Mistral on compare rather than assuming Qwen's budget reputation extends to its higher tier automatically.
See the full current Qwen lineup live above, or find the single cheapest option for your workload on cheapest models.
