AI BriefingGooglePricing & Plans00:00
AI summarized from verified sources
Flex & Priority Inference Tiers for Gemini API
Run background jobs at 50% cost, saving budgets significantly.
SOURCE CHECK
2 sources
Sources
Key Points
- 1Flex: Cost-opt, lower priority
- 2Priority: Low-latency, high priority
- 3For Gemini 2.5/3.1 models
- 4Available now
Google added Flex (cost-optimized) and Priority (latency-optimized) tiers to Gemini API. Flex offers up to 50% savings for tolerant workloads; Priority prioritizes traffic. Balances cost, speed, reliability for devs.
Key point
Google updated the Gemini Developer API pricing page with clearer input, output, and caching rates by model. It also makes free and paid tiers easier to compare for prototype and production planning.
Impact
Separate prototype and production costs more clearly. Key checks: Pricing clarified by model / Input, output, and cache listed / Free vs paid tiers are clearer.