Gemini 2.5 Flash
Released
Gemini 2.5 Flash is an AI model released by Google on Thursday, Apr 17 2025, 23 days after Gemini 2.5 Pro. Benchmark results (shown below) cover BullshitBench v2.
Get Google releases and AI news.
One email, every other Monday.
API pricing
- InputWhat you pay for everything you send the model — your question, plus any documents or earlier conversation you include with it.
- $0.30
- Cached inputA reduced rate for text you send over and over. If every request starts with the same instructions or the same document, the provider keeps a copy ready and charges less to read it again.
- $0.03
- OutputWhat you pay for the text the model writes back. It is normally the dearer half: producing an answer costs more than reading one.
- $2.50
Benchmarks
About
Gemini 2.5 Flash, released April 17, 2025, brought hybrid reasoning to Google's high-volume tier: developers could turn thinking on or off and set an exact token budget for it, trading accuracy against latency and cost per request. It was the first model to expose that dial explicitly, and the design was widely copied.
Positioned as the workhorse of the 2.5 generation, Flash delivered a large share of 2.5 Pro's capability at a fraction of the price, making it one of the most heavily used API models of 2025 for chat products, summarisation, and agent sub-tasks. Flash-Lite followed in June 2025 to cover the ultra-low-cost end of the lineup.
Compare Gemini 2.5 Flash with
Suggested comparisons
Frequently asked questions
Gemini 2.5 Flash was released by Google on Thursday, Apr 17 2025.