Back to Models
google/gemini-3-1-flash-lite-preview
Not Available
Gemini 3.1 Flash Lite Preview
Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities. Improvements span audio input/ASR, RAG snippet ranking, translation, data extraction, and code completion. Supports full thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Priced at half the cost of Gemini 3 Flash.
3/3/2026
1,048,576 tokens
Specifications
Modalities
Input
text
image
video
file
audio
Output
text
Supported Parameters
include_reasoning
max_tokens
reasoning
response_format
seed
stop
structured_outputs
temperature
tool_choice
tools
top_p
Max Output Tokens
65,536Reasoning Configuration
- Default
- Thinking on (effort: minimal)
- Selectable effort levels
- highmediumlowminimal
- Turning thinking off
- Possible (reasoning.enabled: false)
Frequently asked questions
- Is Gemini 3.1 Flash Lite Preview available on FastMetal?
- Not at the moment. GLM 5.3 Flash, from the same lab, is available on the FastMetal API today.
- What is the context window of Gemini 3.1 Flash Lite Preview?
- 1,048,576 tokens, shared between the prompt and the response.