Back to Models
google logoGoogle
gemini-3.7-flash

Gemini 3.7 Flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

8/13/2026
1,048,576 tokens
Input: $0.7913/M
Output: $3.96/M

Specifications

Modalities

Input
text
image
video
file
audio
Output
text

Supported Parameters

include_reasoning
max_tokens
reasoning
reasoning_effort
response_format
seed
stop
structured_outputs
temperature
tool_choice
tools
top_p

Max Output Tokens

65,536

Reasoning Configuration

Default
Thinking on (effort: medium)
Selectable effort levels
high
medium
low
Turning thinking off
Not possible (always thinks)

How to control thinking, and what it costs

Data policy

Prompt retention
Unknown — we could not confirm
Training
Unknown — not declared

"Unknown" does not mean "safe". It means we could not confirm it.

Whether a provider trains on prompts is a declared value from our terms with them. Retention is determined from the upstream listing for every host this model can reach. Neither is guessed.

Measured performance

Time to first response (p50)1.9s
Throughput (p50)516 tokens/s

    Measured through FastMetal's own gateway over the last 30 days (651 first-response samples).

    Throughput is sampled once a day: the median of 26 observations over the last 30 days.

    Time to first response is measured to the first streamed chunk; throughput counts all output tokens, including reasoning tokens. These are observations, not a performance guarantee.

    Code Examples

    curl https://api.fastmetal.ai/v1/chat/completions \
      -H "Content-Type: application/json" \
      -H "Authorization: Bearer YOUR_API_KEY" \
      -d '{
        "model": "gemini-3.7-flash",
        "messages": [{"role": "user", "content": "Hello!"}]
      }'

    How Gemini 3.7 Flash actually answers

    Real responses to our standard prompts, recorded on FastMetal.

    Count the number of 'r's in 'strawberry'

    Count the number of 'r's in 'strawberry'. Explain your reasoning step by step.

    There are **3** 'r's in the word "strawberry". ### Step-by-step reasoning: 1. **Break down the word into individual letters:** * 1st letter: **s** * 2nd letter: **t** * 3rd letter: **r** *(1st 'r')* * 4th letter: **a** * 5th letter: **w** * 6th letter: **b** * 7th letter: **e** * 8th letter: **r** *(2nd 'r')* * 9th letter: **r** *(3rd 'r')* * 10th letter: **y** 2. **Identify the positions of 'r':** * The 3rd letter is an **r** (in "st**r**aw"). * The 8th letter is an **r** (in "be**r**ry"). * The 9th letter is an **r** (in "ber**r**y"). 3. **Total count:** 1 + 1 + 1 = **3**.
    4.5s response530 output tokens

    Debug This Error

    I'm getting the following error in my Node.js application: TypeError: Cannot read properties of undefined (reading 'map') at UserList (/app/components/UserList.js:12:25) at renderWithHooks (/app/node_modules/rea…

    ### Why This Error Happens This error occurs because the `users` prop is `undefined` when the component renders. Since `map` is an Array method, calling `undefined.map()` causes JavaScript to throw a `TypeError`. This usually happens when: 1. Data is being fetched asynchronously (e.g., via `fetch` or `axios`), and the component renders **before** the data has arrived. 2. The parent component forgot to pass the `users` prop (`<UserList />`). 3. The initial state in the parent component was initialized as `undefined` (e.g., `const [users, setUsers] = useState()`). --- ### How to Fix It Here are the best ways to fix it, depending on your use case: #### Solution 1: Default Parameter (Recom…
    9.9s response1355 output tokens

    Code Review

    Please review the following Python function and suggest improvements for readability, performance, and best practices: def get_data(url, retries=3): import requests import time for i in range(retries):…

    Here is a detailed review of your function, followed by refactored versions. --- ### Key Issues & Areas for Improvement 1. **Imports inside the function (PEP 8):** Imports should be placed at the top of the file, not inside a function, to avoid repeated import overhead and improve readability. 2. **Bare `except:` clause (Anti-pattern):** Catching everything with `except:` intercepts `KeyboardInterrupt`, `SystemExit`, and memory errors, making debugging difficult. Catch specific exceptions like `requests.RequestException` and `ValueError` (for JSON decoding). 3. **Retrying on non-retryable errors:** A status code like `404 Not Found` or `401 Unauthorized` will never succeed on a…
    13.2s response2117 output tokens

    Compare these answers side by side with other models →

    Frequently asked questions

    How much does the Gemini 3.7 Flash API cost?
    On FastMetal, Gemini 3.7 Flash is billed per token: $0.7913 per 1M input tokens and $3.96 per 1M output tokens on a US-dollar account, before tax. Usage is drawn from a prepaid balance; there is no subscription or monthly fee.
    Can I call Gemini 3.7 Flash with the OpenAI SDK?
    Yes. Point base_url at https://api.fastmetal.ai/v1 and pass "gemini-3.7-flash" as the model; existing OpenAI-style code works unchanged, including streaming, tool calls and structured output.
    What is the context window of Gemini 3.7 Flash?
    1,048,576 tokens, shared between the prompt and the response.
    What do I need to try Gemini 3.7 Flash?
    Create an account and add credit; the model is then available both in the browser chat and over the API. There is no contract or minimum spend.

    Try Gemini 3.7 Flash right now

    Gemini 3.7 Flash is available on FastMetal through one API key. Start in the browser, or call it from the OpenAI SDK.