Back to Models
google logo
google/gemini-3-1-flash-lite-preview
Not Available

Gemini 3.1 Flash Lite Preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities. Improvements span audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.

3/3/2026
1,048,576 tokens

Specifications

Modalities

Input
text
image
video
file
audio
Output
text

Supported Parameters

include_reasoning
max_tokens
reasoning
response_format
seed
stop
structured_outputs
temperature
tool_choice
tools
top_p

Max Output Tokens

65,536

Reasoning Configuration

Default
Thinking on (effort: minimal)
Selectable effort levels
high
medium
low
minimal
Turning thinking off
Possible (reasoning.enabled: false)

How to control thinking, and what it costs

Frequently asked questions

Is Gemini 3.1 Flash Lite Preview available on FastMetal?
Not at the moment. GLM 5.3 Flash, from the same lab, is available on the FastMetal API today.
What is the context window of Gemini 3.1 Flash Lite Preview?
1,048,576 tokens, shared between the prompt and the response.