Pricing Verified fallback snapshot · 43 models · 7 providers · verified September 16, 2026.
Google Gemini pricing

Gemini API Pricing & Model Costs

Compare Gemini API input, output and cached-token pricing across Flash, Pro and preview models, then run the same workload through TokenCOGS calculators.

Gemini models tracked10
Lowest input rate$0.1Gemini 2.5 Flash-Lite
Lowest output rate$0.4Gemini 2.5 Flash-Lite
Largest tracked context1.048576M tokensGemini 3.8 Flash
At a glance

Popular Gemini API price points

Start with the models developers are most likely to compare, then use the full table for the rest of the Gemini lineup.

Current rates

Gemini API pricing table

USD per 1M text tokens for Google Gemini models. Compare standard input, output and cached-token rates before modeling your workload.

Pricing methodology →
ModelInputOutputCached inputContext
Gemini 3.8 Flash$0.75$3.75$0.0751.048576M tokensDetails →
Gemini 3.7 FlashCurrent / previous generation$0.75$3.75$0.0751M tokensDetails →
Gemini 3.6 Flash$0.75$3.75$0.0751M tokensCalculate →
Gemini 3.5 Flash$1.5$9$0.15Not listed in registryCalculate →
Gemini 3.5 Flash-Lite$0.3$2.5$0.03Not listed in registryCalculate →
Gemini 3.1 Pro Preview Preview$2$12$0.21M tokensDetails →
Gemini 3.1 Flash-Lite$0.25$1.5$0.025Not listed in registryCalculate →
Gemini 2.5 Pro$1.25$10$0.1251M tokensDetails →
Gemini 2.5 Flash$0.3$2.5$0.031M tokensCalculate →
Gemini 2.5 Flash-Lite$0.1$0.4$0.01Not listed in registryCalculate →
Deals & credits

Current Google savings programs

Current promotion

50% promotional credit

50% billing credit on Gemini 3.6 / 3.7 / 3.8 Flash

See details ↗
Credits program

Up to $350K in Cloud credits

Google for Startups AI cloud credits

See details ↗
How to choose

Price is only the first filter

Use this table to eliminate models that do not fit your unit economics, then benchmark the remaining candidates on your own task for output quality, latency, rate limits and operational reliability.

Gemini 2.5 Flash-Lite currently has the lowest tracked input-token price in this Google set. That does not automatically make it the lowest-cost model for every workload because output ratios, caching, long-context rules and tool usage can change the result.