← All providers

Z.AI API pricing

6 Z.AI models tracked, updated daily. Full rates below, with guidance on which tier fits which workload.

ModelTierInputOutputContext
GLM 4.7 FlashBudget$0.06$0.40203K
GLM 5.2Budget$0.75$2.351M
GLM 5Budget$0.95$2.55205K
GLM 5.1Budget$0.97$3.04205K
GLM 5 TurboStandard$1.20$4.00203K
GLM 5V TurboStandard$1.20$4.00203K

Cached input rates and full detail on the pricing page. Prices in USD per 1M tokens.

Z.AI (formerly Zhipu AI) is one of the few providers where most of the current lineup is genuinely open-weight — the GLM line ships its weights publicly under a permissive license, with a smaller set of "turbo" variants held back as closed, API-only tiers. See the live table above for current rates, and the open-vs-closed race to see exactly how competitive Z.AI's open releases are against closed frontier models.

Open by default, "turbo" is the exception

The numbered GLM releases (GLM-4.5, GLM-4.6, GLM-4.7, GLM-5, GLM-5.1, GLM-5.2 and their variants) are open-weight — downloadable and self-hostable, which is unusual among providers competing at the frontier. The "-turbo" suffixed variants (e.g. a GLM-5-turbo or GLM-5v-turbo) are the closed, API-only counterparts, typically optimized for latency or cost rather than being materially different models. If self-hosting or license flexibility matters to you, confirm you're looking at a non-turbo release before assuming it's open.

Why this matters for pricing

Because most of the lineup is open-weight, Z.AI's own API pricing tends to be aggressive — it's competing against the possibility that a customer simply self-hosts the same weights elsewhere. This is a useful signal in general: providers whose flagship models are also freely downloadable usually can't sustain a large API markup, since price-sensitive customers have a credible alternative.

Practical guidance

  • Default to the latest numbered GLM release for general chat and coding — the line has been competitive on both LiveBench and coding-specific benchmarks like SWE-bench Verified.
  • Check vision/tool-calling variants separately (GLM releases with a "V" designation add multimodal input) if your workload needs image understanding — these are typically separate SKUs from the text-only line.
  • Don't assume "turbo" means better — it usually means a closed, latency-optimized variant, not a capability upgrade.

When Z.AI vs. a competitor

For coding and general-purpose work, compare GLM directly against DeepSeek, Qwen, and Moonshot — the cluster of providers where open-weight, competitively-priced models are the norm rather than the exception. See how GLM's coding benchmarks stack up on SWE-bench Verified and Aider Polyglot, or put it head-to-head with any model on compare.

See current GLM pricing live above, or find where it ranks on cheapest models.

Frequently asked questions

How many Z.AI models does TokenCost track?
TokenCost currently tracks 6 Z.AI models, refreshed daily from live pricing data.
What is the cheapest Z.AI model?
GLM 4.7 Flash is currently Z.AI’s lowest-cost model on input price, at $0.06 per 1M input tokens and $0.40 per 1M output tokens.
Does Z.AI support prompt caching?
Where a cached-input rate is published, TokenCost shows it directly in the pricing table (cached input is typically billed at a steep discount versus the standard input rate). See our prompt caching guide for the general mechanics.