Skip to content
  • Models
  • Rankings
  • Ori
Sign Up
Sign Up
OpenRouterOpenRouter
© 2026 OpenRouter, Inc

Product

  • Chat
  • Rankings
  • Benchmarks
  • Apps
  • Discover
  • Models
  • Collections
  • Providers
  • Pricing
  • Business
  • Enterprise
  • Labs

Company

  • About
  • Blog
  • Careers
    Hiring
  • Privacy
  • Terms of Service
  • Trust Center
  • Support
  • Works With OR
  • Data
  • Brand

Developer

  • Documentation
  • API Reference
  • Developer Platform
  • Status

Connect

  • Discord
  • GitHub
  • LinkedIn
  • X
  • YouTube

Model usage data, product updates, and research reports. One email each week.

By subscribing you agree to receive the OpenRouter newsletter: model usage data, product updates, and research reports, about one email a week. Unsubscribe anytime via the link in every email. See our Privacy Policy.

Favicon for google

Google: Gemini 2.5 Flash

Going away October 20, 2026

google/gemini-2.5-flash

Compare

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater accuracy and nuanced context handling.

Additionally, Gemini 2.5 Flash is configurable through the "max tokens for reasoning" parameter, as described in the documentation (https://openrouter.ai/docs/use-cases/reasoning-tokens#max-tokens-for-reasoningOpens in new tab).

Modalities

In / Out Price

$0.30 / $2.50per 1M

Context

1.0M

Released

Jun 17, 2025

Knowledge Cutoff

Jan 2025

Compare
ProvidersPricingPerformanceUptimeBenchmarksAppsActivityFAQExplore

Providers

Different companies host the same model. OpenRouter routes your request to one of them based on the routing mode you pick — Balanced (price + speed), Nitro (fastest), Floor (cheapest), or Exacto (highest tool-calling accuracy).

Pricing

The average price customers actually pay for this model, next to the prices providers post. Caching and discounts mean the price actually paid is often well below the listed one.

Performance

Throughput is how fast the model writes (tokens per second — higher is better). Latency is total round-trip time (lower is better). TTFT is time-to-first-token — how long before you see anything appear (lower is better).

Uptime

Uptime is the percentage of the past 3 days that at least one provider was responding to requests. Availability is the percentage of time that inference was successfully served. OpenRouter continuously monitors and uses the next-best provider when one returns an error.

Benchmarks

Scores on standardized evaluations. Higher percentages are better — and rank percentile shows where this model lands among all models on OpenRouter.

Benchmark score summary for Google: Gemini 2.5 Flash (Artificial Analysis and Design Arena)
SourceBenchmarkScore
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) GPQA Diamond68.3%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) HLE4.7%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) IFBench39.0%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) τ²-Bench Telecom14.9%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) AA-LCR49.9%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) CritPt1.4%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) Terminal-Bench Hard12.1%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) AA-Omniscience Accuracy26.1%
Artificial AnalysisGemini 2.5 Flash (Non-reasoning) AA-Omniscience Non-Hallucination Rate7.0%
Artificial AnalysisGemini 2.5 Flash Preview (Non-reasoning) GPQA Diamond59.4%
Artificial AnalysisGemini 2.5 Flash Preview (Non-reasoning) HLE3.8%
Artificial AnalysisGemini 2.5 Flash (Reasoning) GPQA Diamond79.0%
Artificial AnalysisGemini 2.5 Flash (Reasoning) HLE12.1%
Artificial AnalysisGemini 2.5 Flash (Reasoning) IFBench50.3%
Artificial AnalysisGemini 2.5 Flash (Reasoning) τ²-Bench Telecom31.6%
Artificial AnalysisGemini 2.5 Flash (Reasoning) AA-LCR65.3%
Artificial AnalysisGemini 2.5 Flash (Reasoning) CritPt1.1%
Artificial AnalysisGemini 2.5 Flash (Reasoning) Terminal-Bench Hard13.6%
Artificial AnalysisGemini 2.5 Flash (Reasoning) AA-Omniscience Accuracy26.0%
Artificial AnalysisGemini 2.5 Flash (Reasoning) AA-Omniscience Non-Hallucination Rate24.7%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) GPQA Diamond76.6%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) HLE8.7%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) IFBench43.5%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) τ²-Bench Telecom28.4%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) AA-LCR60.0%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) CritPt0.0%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) Terminal-Bench Hard14.4%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) AA-Omniscience Accuracy26.8%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Non-reasoning) AA-Omniscience Non-Hallucination Rate8.8%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) GPQA Diamond79.3%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) HLE13.8%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) IFBench52.3%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) τ²-Bench Telecom45.6%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) AA-LCR71.0%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) CritPt0.3%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) Terminal-Bench Hard16.7%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) AA-Omniscience Accuracy28.3%
Artificial AnalysisGemini 2.5 Flash Preview (Sep '25) (Reasoning) AA-Omniscience Non-Hallucination Rate10.1%
Artificial AnalysisGemini 2.5 Flash Preview (Reasoning) GPQA Diamond69.8%
Artificial AnalysisGemini 2.5 Flash Preview (Reasoning) HLE12.1%
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena 3D Elo1100
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena Code Categories Elo1119
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena Data Visualization Elo1149
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena Game Development Elo1089
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena SVG Elo1043
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena UI Component Elo1108
Design ArenaGemini 2.5 Flash Preview 09-2025 Models Arena Website Elo1127

Apps

Public apps that send the most traffic to this model. Good signal for what real production workloads look like — and a hint at which use cases this model is best suited for.

Activity

Token volume and request traffic to this model over time.

Quick Start

Drop-in code to call this model. OpenRouter's API is OpenAI-compatible — most SDKs work by just swapping the base URL. The only thing that changes between models is the model slug below.

Explore more models

AI Models with Vision: Multimodal LLMs for Image UnderstandingCollectionAI Model RankingsRanking
$0.30$2.50$0.03$0.100.55s92 tps
99.45%
$0.30$2.50$0.03$0.100.81s69 tps
99.47%
$0.30$2.50$0.03$0.100.57s91 tps
99.97%
$0.30$2.50$0.03$0.101.27s85 tps
85.80%
Not used in Standard routing:Why these endpoints are not used
Priority
$0.54$4.50$0.054$0.180.74s93 tps
99.98%
Flex
$0.15$1.25$0.015$0.052.55s4 tps
100.00%
Priority
$0.54$4.50$0.054$0.180.40s35 tps
100.00%

Throughput

93tok/s

P50, best across providers

Latency

0.40s

P50, best provider

AutoExacto Benchmarks
GPQA DiamondTAU-BenchGoogle Vertex (EU)74.9%58.7%Google Vertex73.6%59.3%Google Vertex (Global)73.2%59.3%Google AI Studio76.3%54.7%auto-routing73.2%57.3%
Uptime (3d)The model was reachable. Request routed to a provider.

100.00%

Availability (3d)The model returned inference from any provider. Errors and empty responses count against it.

99.83%

Availability over the last 3 days

Last 72 hours
Availability 99.83%
3 Days Ago2 Days AgoYesterdayNow

Availability over the last 24 hours

OpenRouter Availability
99.78%
Without Routing
98.35%

When an error occurs in an upstream provider, we can recover by routing to another healthy provider, if your request filters allow it. You can access per-provider uptime data programmatically through the Endpoints API. Learn more about our load balancing and customization options.

1.
Favicon for https://docs-api.payready.com/
payready-document-extraction
new
68.3Btokens
2.
Favicon for https://i.love.koalas.ai/
KoalaBear
new
66.5Btokens
3.
Favicon for https://pinxai.com/
PinxAi
new
47.3Btokens
4.
Favicon for https://nousresearch.com
Hermes Agent
Hermes Agent is an open-source, self-improving AI agent by Nous Research that runs persistently with memory across sessions, and builds reusable skills from experience. It comes with 40+ built-in tools, including web search, browser automation, and vision, plus scheduled automations and subagents.
39.9Btokens
5.
Favicon for https://openclaw.ai/
OpenClaw
OpenClaw is an open-source AI agent that connects to your messaging apps and takes real actions on your behalf, from running commands and browsing the web to managing files and sending emails.
30.4Btokens

Frequently asked questions

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater accuracy and nuanced context handling.

Gemini 2.5 Flash costs $0.30/M input tokens and $2.50/M output tokens, with separate rates for Cache Read at $0.03/M tokens, Cache Write at $0.08333/M tokens, Image Input at $0.30/M tokens, Input Audio at $1.00/M tokens, and Input Audio Cache at $0.10/M tokens.

Gemini 2.5 Flash has a 1,048,576 token context window. It supports up to 65,535 completion tokens.

Yes. Gemini 2.5 Flash accepts tools and tool_choice for function calling. It also supports structured outputs via a JSON schema in response_format.

Gemini 2.5 Flash accepts files such as PDFs, images, text, audio, and video as input and returns text.

Gemini 2.5 Flash is served by 2 providers on OpenRouter: Google Vertex and Google AI Studio. Requests are routed to the best available provider, with automatic failover to the others, and you can pin or exclude providers with provider routing.

Gemini 2.5 Flash was released on June 17, 2025. Its knowledge cutoff is January 31, 2025.

More models from Google

Gemini 3.8 Flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Text1.0M context$0.75 / $3.75
Gemini 3.8 Flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Text1.0M context$0.375 / $1.875
Gemini 3.7 Flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

Text1.0M context$0.75 / $3.75
Gemini 3.7 Flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step problem solving.

Text1.0M context$0.375 / $1.875
Gemini 3.6 Flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Text1.0M context$0.75 / $3.75
Gemini 3.6 Flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and less hedging, while reducing token use and the number of model calls needed to complete a task.

Text1.0M context$0.375 / $1.875
Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Text1.0M context$0.30 / $2.50
Gemini 3.5 Flash Lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Text1.0M context$0.15 / $1.25
Nano Banana 2 Lite

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation in roughly 4 seconds — about 2.7× faster than Gemini 3.1 Flash Image — while keeping the character consistency, precise editing, and real-world knowledge of the Nano Banana family.

A single drop-in API handles text-to-image, image editing, and multi-image composition. As a multimodal model it also returns text alongside images. Outputs are generated at 1K resolution across 14 aspect ratios and carry an invisible SynthID watermark so they can be identified as AI-generated.

Positioned as the best balance of quality and speed in the Nano Banana 2 line, it lets you generate thousands of images at a fraction of the cost of heavier production models — ideal for prototyping, real-time apps, and visual workflows at scale.

Image66K context$0.25 / $30
Nano Banana 2

Gemini 3.1 Flash Image, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed. It combines advanced contextual understanding with fast, cost-efficient inference, making complex image generation and iterative edits significantly more accessible. Aspect ratios can be controlled with the image_config API Parameter

Image131K context$0.50 / $60
Nano Banana Pro

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthesis. The model generates context-rich graphics, from infographics and diagrams to cinematic composites, and can incorporate real-time information via Search grounding.

It offers industry-leading text rendering in images (including long passages and multilingual layouts), consistent multi-image blending, and accurate identity preservation across up to five subjects. Nano Banana Pro adds fine-grained creative controls such as localized edits, lighting and focus adjustments, camera transformations, and support for 2K/4K outputs and flexible aspect ratios. It is designed for professional-grade design, product visualization, storyboarding, and complex multi-element compositions while remaining efficient for general image creation workflows.

Image131K context$2 / $120
Gemini Embedding 2

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.

Embeddings$0.20/M tokens
Gemini Embedding 2

Gemini Embedding 2 is Google's first multimodal embedding model. We currently support mapping text and images into a unified vector space for semantic search and retrieval-augmented generation (RAG). It supports input context up to 8,192 tokens and flexible output dimensions from 128 to 3,072 (recommended: 768, 1536, or 3,072). Designed for cross-modal similarity — you can embed a text query and retrieve the most relevant images, or vice versa — making it well-suited for multimodal search, recommendation, and document understanding pipelines.

Embeddings$0.10/M tokens
Gemini 3.5 Flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution loops, supporting text, image, video, audio, and PDF inputs.

Defaults to medium thinking effort for faster and more cost-efficient responses, with full support for thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs.

Text1.0M context$1.50 / $9
Gemini 3.5 Flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution loops, supporting text, image, video, audio, and PDF inputs.

Defaults to medium thinking effort for faster and more cost-efficient responses, with full support for thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs.

Text1.0M context$0.75 / $4.50
Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints.

Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.

Text1.0M context$0.25 / $1.50
Gemini 3.1 Flash Lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic workflows, simple data extraction, and applications where responsiveness and API cost are the primary constraints.

Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.

Text1.0M context$0.125 / $0.75
Chirp 3

Chirp 3 is Google's latest multilingual speech-to-text model. It offers enhanced transcription accuracy across 24 GA languages and 77+ preview languages, with support for automatic language detection, automatic punctuation, and a built-in denoiser for cleaner audio processing.

Transcription$0.016/minute
Gemini Pro Latest

This model always redirects to the latest model in the Gemini Pro family.

Text1.0M context
Gemini Flash Latest

This model always redirects to the latest model in the Gemini Flash family.

Text1.0M context