Back to Models
google logo
google/gemini-2-5-flash-lite
Not Available

Gemini 2.5 Flash Lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance across common benchmarks compared to earlier Flash models. By default, "thinking" (i.e. multi-pass reasoning) is disabled to prioritize speed, but developers can enable it via the [Reasoning API parameter](https://openrouter.ai/docs/use-cases/reasoning-tokens) to selectively trade off cost for intelligence.

7/22/2025
1,048,576 tokens

Specifications

Modalities

Input
text
image
file
audio
video
Output
text

Supported Parameters

include_reasoning
max_tokens
reasoning
response_format
seed
stop
structured_outputs
temperature
tool_choice
tools
top_p

Max Output Tokens

65,535

Frequently asked questions

Is Gemini 2.5 Flash Lite available on FastMetal?
Not at the moment. GLM 5.3 Flash, from the same lab, is available on the FastMetal API today.
What is the context window of Gemini 2.5 Flash Lite?
1,048,576 tokens, shared between the prompt and the response.