Google Gemini API: models and pricing

FastMetal serves 8 Google Gemini models through one OpenAI-compatible API key, billed per token from a prepaid balance. Change the base URL and key, and existing code calls them as is. The list also shows Google Gemini models we do not serve yet.

Sort by:

41 models

Gemini 3.8 Flash

gemini-3.8-flash
Policy unknown

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Provider:google logogoogle
Pricing:$0.7913 in · $3.96 out / 1M tokens
Context:1.0M tokens

Gemini 3.7 Flash

gemini-3.7-flash
Policy unknown

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Provider:google logogoogle
Pricing:$0.7913 in · $3.96 out / 1M tokens
Context:1.0M tokens

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#71

Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

google/gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:66K tokens

Nano Banana 2 (Gemini 3.1 Flash Image)

google-nano-banana-2

Google's fast, high-quality image generation model.

Provider:google logogoogle
Pricing:$0.078 per image
Context:131K tokens

Nano Banana Pro (Gemini 3 Pro Image)

google/gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens

Gemini 3.5 Flash

gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Provider:google logogoogle
Pricing:$1.65 in · $9.90 out / 1M tokens
Context:1.0M tokens

Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Google Gemini Pro Latest

~google/gemini-pro-latest

This model always redirects to the latest model in the Google Gemini Pro family.

Provider:~google
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemma 4 26B A4B

gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Provider:google logogoogle
Pricing:$0.0739 in · $0.3587 out / 1M tokens
Context:262K tokens

Gemma 4 26B A4B (free)

google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:262K tokens

Gemma 4 31B

japan-gemma-4-31b

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Provider:google logogoogle
Pricing:$0.1657 in · $0.6436 out / 1M tokens
Context:262K tokens
Text Overall
#76

Gemma 4 31B (free)

google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:262K tokens

Gemini 3.1 Flash Lite Preview

google/gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#109

Nano Banana 2 (Gemini 3.1 Flash Image Preview)

google/gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:66K tokens

Gemini 3.1 Pro Preview Custom Tools

google/gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party or user-defined functions are available.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemini 3.1 Pro Preview

gemini-3.1-pro-preview
Policy unknown

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows.

Provider:google logogoogle
Pricing:$2.11 in · $12.66 out / 1M tokens
Context:1.0M tokens
Text Overall
#18

Gemini 3 Flash Preview

google/gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Nano Banana Pro (Gemini 3 Pro Image Preview)

google/gemini-3-pro-image-preview

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthe…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:66K tokens

Gemini 3 Pro Preview

google/gemini-3-pro-preview

Gemini 3 Pro is Google’s flagship frontier model for high-precision multimodal reasoning, combining strong performance across text, image, video, audio, and code with a 1M-token context window.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Nano Banana (Gemini 2.5 Flash Image)

google/gemini-2.5-flash-image

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens

Gemini 2.5 Flash Lite Preview 09-2025

google/gemini-2.5-flash-lite-preview-09-2025

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemini 2.5 Flash Lite

google/gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemini 2.5 Flash

google/gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#148

Gemini 2.5 Pro

google/gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#86

Gemini 2.5 Pro Preview 06-05

google/gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemma 3n 4B

google/gemma-3n-e4b-it

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#264

Gemma 3n 4B (free)

google/gemma-3n-e4b-it:free

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:8K tokens

Gemini 2.5 Pro Preview 05-06

google/gemini-2.5-pro-preview-05-06

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemma 3 12B

google/gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#230

Gemma 3 12B (free)

google/gemma-3-12b-it:free

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens

Gemma 3 4B

google/gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#281

Gemma 3 4B (free)

google/gemma-3-4b-it:free

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens

Gemma 3 27B

google/gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#197

Gemma 3 27B (free)

google/gemma-3-27b-it:free

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens

Gemini 2.0 Flash Lite

google/gemini-2.0-flash-lite-001

Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5), all…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemini 2.0 Flash

google/gemini-2.0-flash-001

Gemini Flash 2.0 offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5).

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#204

Gemma 2 27B

google/gemma-2-27b-it

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini).

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#289

Gemma 2 9B

google/gemma-2-9b-it

Gemma 2 9B by Google is an advanced, open-source language model that sets a new standard for efficiency and performance in its size class.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#310

Gemini Flash Lite

gemini-flash-lite-free
Trains on prompts

Google's fastest and most cost-efficient model in the Gemini series. Delivers frontier-class performance with 2.5x faster time-to-first-token, ideal for high-volume, latency-sensitive applications.

Provider:google logogoogle
Pricing:$0.00 in · $0.00 out / 1M tokens
Context:1.0M tokens