Google API pricing
14 Google models tracked, updated daily. Full rates below, with guidance on which tier fits which workload.
Cached input rates and full detail on the pricing page. Prices in USD per 1M tokens.
Google's lineup spans the widest price range of any major provider we track — from the genuinely low-cost, open-weight Gemma models up through Flash Lite, Flash, and the Pro-tier Gemini models. That range makes tier selection especially important here: the gap between the cheapest and most capable Google model is large, and picking the wrong end for your task wastes either money or capability. See the live table above for current rates.
Gemma vs. Gemini — a real distinction, not just a naming difference
The Gemma models are Google's open-weight family and sit in the Budget tier alongside Gemini's Flash Lite variants. They're a reasonable choice for high-volume, low-complexity tasks, and — because they're open-weight — some teams also self-host them for workloads with very high, predictable volume where API pricing no longer makes sense at all (see Open-Source vs API for when that trade-off tips in favor of self-hosting). Gemini proper (Flash, Flash Lite, and Pro) is the hosted API line and where most teams start.
Flash vs. Pro
Flash and Flash Lite are built for high-throughput, cost-sensitive workloads — classification, summarization, chat at scale — while the Pro-tier models (including the "Custom Tools" variant tuned for agentic tool-calling) are positioned for harder reasoning and multi-step tasks. The practical guidance mirrors every provider's tiering: default to Flash, and only move to Pro for the subset of requests where you have evidence Flash's capability ceiling is the limiting factor.
Long-context costs deserve special attention on Gemini
Google's Gemini models are frequently used specifically for their large context windows — but a big context window is a capability, not a discount. Every token you send into that window is billed at the standard input rate regardless of how much of the window you're using. If you're using Gemini for a very-large-context use case (a large document, a big codebase), make sure you're not paying to send context the model doesn't actually need — see Which Frontier Model Is Actually Cheapest for Your Workload? for how input-heavy workloads should weight the comparison, and How to Cut Your LLM API Bill by 80% for context-trimming tactics that apply directly here.
The free tier and rate limits
Google has historically offered a free usage tier for Gemini models with meaningful rate limits, which makes it a common starting point for prototypes and side projects before a workload graduates to paid, higher-throughput usage. Confirm current free-tier terms directly with Google before relying on them for anything production-facing, since free-tier availability and limits change independently of the paid pricing shown here.
When Google vs. a competitor
Flash-tier Gemini models are consistently competitive on price for general-purpose and high-volume workloads, making Google a strong default to check against OpenAI and Anthropic before committing elsewhere. For pure lowest-cost budget-tier work, also compare against DeepSeek and Qwen, which frequently undercut even Gemini's Flash Lite tier.
See current Gemma-through-Pro pricing live above, or run Google against any competitor on compare.
