Models

Every model here is reachable through the same OpenAI-compatible endpoint with one API key. Prices are per million tokens in yen, billed from your prepaid balance; models marked "not on FastMetal yet" are listed for reference and cannot be called.

Sort by:

303 models

DeepSeek V4.1 Flash

deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Provider:deepseek logodeepseek
Pricing:¥50.686 in · ¥202.7441 out / 1M tokens
Context:1.0M tokens

Mercury 2.5

inception/mercury-2.5

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...

Provider:inception
Pricing:¥6.768 in · ¥25.3799 out / 1M tokens
Context:260K tokens

GPT-6 Astra

openai/gpt-6-astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...

Provider:openai logoopenai
Pricing:¥1,717.723 in · ¥8,588.615 out / 1M tokens
Context:1.1M tokens

GPT-6 Astra Pro

openai/gpt-6-astra-pro

GPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.

Provider:openai logoopenai
Pricing:¥1,717.723 in · ¥8,588.615 out / 1M tokens
Context:1.1M tokens

Qwen3.8 Max (0902)

qwen/qwen3.8-max-0902

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Provider:qwen logoqwen
Pricing:¥357.4 in · ¥1,072.2 out / 1M tokens
Context:1.0M tokens

Gemini 3.8 Flash

google/gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.

Provider:google logogoogle
Pricing:¥128.8292 in · ¥644.1461 out / 1M tokens
Context:1.0M tokens

Muse Spark 1.3

meta/muse-spark-1.3

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...

Provider:meta logometa
Pricing:¥219.8968 in · ¥747.6489 out / 1M tokens
Context:1.0M tokens

Muse Spark 1.3 Contributor

meta/muse-spark-1.3-contributor

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...

Provider:meta logometa
Pricing:Not on FastMetal yet
Context:1.0M tokens

Claude Fable 5.1

anthropic/claude-fable-5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...

Provider:anthropic logoanthropic
Pricing:¥1,787 in · ¥8,935 out / 1M tokens
Context:1.0M tokens

Hy4 preview

tencent/hy4-preview

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...

Provider:tencent logotencent
Pricing:Not on FastMetal yet
Context:1.0M tokens

GLM 5.3 Flash

z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...

Provider:z-ai logoz-ai
Pricing:¥25.343 in · ¥84.4767 out / 1M tokens
Context:1.3M tokens
Text Overall
#29

Qwen3.8 Flash

qwen/qwen3.8-flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens

Muse Spark 1.2 Contributor

meta/muse-spark-1.2-contributor

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...

Provider:meta logometa
Pricing:Not on FastMetal yet
Context:1.0M tokens

GLM 5.3

z-ai/glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...

Provider:z-ai logoz-ai
Pricing:¥250.18 in · ¥786.28 out / 1M tokens
Context:1.0M tokens

Qwen3.8 27B

qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Provider:qwen logoqwen
Pricing:¥80.415 in · ¥571.84 out / 1M tokens
Context:1.0M tokens
Text Overall
#89

Gemini 3.7 Flash

google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...

Provider:google logogoogle
Pricing:¥134.025 in · ¥670.125 out / 1M tokens
Context:1.0M tokens

Grok 4.6

x-ai/grok-4.6

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Provider:x-ai logox-ai
Pricing:¥357.4 in · ¥1,072.2 out / 1M tokens
Context:500K tokens

Qwen3.8 2.4T A95B

qwen/qwen3.8-2.4t-a95b

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Provider:qwen logoqwen
Pricing:¥351.8348 in · ¥1,055.5044 out / 1M tokens
Context:1.0M tokens

Muse Glimmer 30B

meta/muse-glimmer-30b

Meta's dense, open-weight 30B multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for long-horizon autonomous-agent tasks on consumer hardware.

Provider:meta logometa
Pricing:¥62.545 in · ¥268.05 out / 1M tokens
Context:131K tokens

Solar Pro 4

upstage/solar-pro4

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...

Provider:upstage
Pricing:¥15.2222 in · ¥60.8888 out / 1M tokens
Context:524K tokens
Text Overall
#175

Muse Spark 1.2

meta/muse-spark-1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...

Provider:meta logometa
Pricing:¥219.8968 in · ¥747.6489 out / 1M tokens
Context:1.0M tokens

Qwen3.8 Max

qwen/qwen3.8-max

Alibaba's flagship Qwen3.8 model: a 2.4T-parameter MoE (95B active) multimodal reasoning model with a 1M-token context window.

Provider:qwen logoqwen
Pricing:¥357.4 in · ¥1,072.2 out / 1M tokens
Context:1.0M tokens
Text Overall
#22

Claude Opus 5

anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...

Provider:anthropic logoanthropic
Pricing:¥893.5 in · ¥4,467.5 out / 1M tokens
Context:1.0M tokens

Gemini 3.5 Flash Lite

google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#57

Gemini 3.6 Flash

google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Inkling

thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...

Provider:thinkingmachines
Pricing:¥178.7 in · ¥723.735 out / 1M tokens
Context:524K tokens
Text Overall
#83

Kimi K3

moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Provider:moonshotai logomoonshotai
Pricing:¥536.1 in · ¥2,680.5 out / 1M tokens
Context:1.0M tokens

Muse Spark 1.1

meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...

Provider:meta logometa
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#10

GPT-5.6 Luna

openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...

Provider:openai logoopenai
Pricing:¥35.74 in · ¥214.44 out / 1M tokens
Context:1.1M tokens

GPT-5.6 Sol

openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...

Provider:openai logoopenai
Pricing:¥893.5 in · ¥5,361 out / 1M tokens
Context:1.1M tokens

GPT-5.6 Terra

openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...

Provider:openai logoopenai
Pricing:¥357.4 in · ¥2,144.4 out / 1M tokens
Context:1.1M tokens

Grok 4.5

x-ai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Provider:x-ai logox-ai
Pricing:¥357.4 in · ¥1,072.2 out / 1M tokens
Context:500K tokens
Text Overall
#38

Hy3

tencent/hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

Provider:tencent logotencent
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#60

Laguna XS 2.1

poolside/laguna-xs-2.1

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Provider:poolside logopoolside
Pricing:Not on FastMetal yet
Context:262K tokens

Laguna XS 2.1 (free)

poolside/laguna-xs-2.1:free

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

Provider:poolside logopoolside
Pricing:Not on FastMetal yet
Context:262K tokens

Claude Sonnet 5

anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

Provider:anthropic logoanthropic
Pricing:¥357.4 in · ¥1,787 out / 1M tokens
Context:1.0M tokens

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)

google/gemini-3.1-flash-lite-image

Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:66K tokens

Nano Banana 2 (Gemini 3.1 Flash Image)

google/gemini-3.1-flash-image

Google's fast, high-quality image generation model.

Provider:google logogoogle
Pricing:¥12.23 per image
Context:131K tokens

Nano Banana Pro (Gemini 3 Pro Image)

google/gemini-3-pro-image

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens

GLM 5.2

z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

Provider:z-ai logoz-ai
Pricing:¥250.18 in · ¥786.28 out / 1M tokens
Context:1.0M tokens

Fusion

openrouter/fusion

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

Provider:openrouter logoopenrouter
Pricing:Not on FastMetal yet
Context:1.0M tokens

Kimi K2.7 Code

moonshotai/kimi-k2.7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...

Provider:moonshotai logomoonshotai
Pricing:¥55 in · ¥530 out / 1M tokens
Context:262K tokens

Claude Fable 5

anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

Provider:anthropic logoanthropic
Pricing:¥1,787 in · ¥8,935 out / 1M tokens
Context:1.0M tokens
Text Overall
#1

Nemotron 3 Ultra

nvidia/nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

Provider:nvidia logonvidia
Pricing:Not on FastMetal yet
Context:512K tokens

Qwen3.7 Plus

qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#61

MiniMax M3

minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

Provider:minimax logominimax
Pricing:¥53.61 in · ¥214.44 out / 1M tokens
Context:1.0M tokens
Text Overall
#80

Step 3.7 Flash

stepfun/step-3.7-flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...

Provider:stepfun logostepfun
Pricing:Not on FastMetal yet
Context:262K tokens

Claude Opus 4.8

anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

Provider:anthropic logoanthropic
Pricing:¥893.5 in · ¥4,467.5 out / 1M tokens
Context:1.0M tokens
Text Overall
#33

Claude Opus 4.8 (Fast)

anthropic/claude-opus-4.8-fast

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3.7 Max

qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Provider:qwen logoqwen
Pricing:¥263.5825 in · ¥790.7475 out / 1M tokens
Context:1.0M tokens

Gemini 3.5 Flash

google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

Provider:google logogoogle
Pricing:¥268.05 in · ¥1,608.3 out / 1M tokens
Context:1.0M tokens

Claude Opus 4.7 (Fast)

anthropic/claude-opus-4.7-fast

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:1.0M tokens

Gemini 3.1 Flash Lite

google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Granite 4.1 8B

ibm-granite/granite-4.1-8b

Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window and is designed for enterprise tasks...

Provider:ibm-granite
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#265

Grok 4.3

x-ai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#79

Mistral Medium 3.5

mistralai/mistral-medium-3-5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#103

Laguna M.1

poolside/laguna-m.1

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...

Provider:poolside logopoolside
Pricing:Not on FastMetal yet
Context:262K tokens

Laguna M.1 (free)

poolside/laguna-m.1:free

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...

Provider:poolside logopoolside
Pricing:Not on FastMetal yet
Context:262K tokens

Google Gemini Pro Latest

~google/gemini-pro-latest

This model always redirects to the latest model in the Google Gemini Pro family.

Provider:~google
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3.6 Max Preview

qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#51

DeepSeek V4 Flash

deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Provider:deepseek logodeepseek
Pricing:¥22.8496 in · ¥49.2146 out / 1M tokens
Context:1.0M tokens
Text Overall
#91

DeepSeek V4 Pro

deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Provider:deepseek logodeepseek
Pricing:¥341.317 in · ¥684.421 out / 1M tokens
Context:1.0M tokens
Text Overall
#55

GPT-5.5

openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.1M tokens
Text Overall
#24

GPT-5.5 Pro

openai/gpt-5.5-pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.1M tokens

Hy3 preview

tencent/hy3-preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...

Provider:tencent logotencent
Pricing:Not on FastMetal yet
Context:262K tokens

MiMo-V2.5

xiaomi/mimo-v2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...

Provider:xiaomi
Pricing:¥25.018 in · ¥50.036 out / 1M tokens
Context:1.1M tokens
Text Overall
#94

MiMo-V2.5-Pro

xiaomi/mimo-v2.5-pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....

Provider:xiaomi
Pricing:¥77.7345 in · ¥155.469 out / 1M tokens
Context:1.1M tokens
Text Overall
#41

GPT-5.4 Image 2

openai/gpt-5.4-image-2

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:272K tokens

Kimi K2.6

moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...

Provider:moonshotai logomoonshotai
Pricing:¥169.765 in · ¥714.8 out / 1M tokens
Context:262K tokens
Text Overall
#50

Claude Opus 4.7

anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

Provider:anthropic logoanthropic
Pricing:¥840 in · ¥4,200 out / 1M tokens
Context:1.0M tokens
Text Overall
#7

GLM 5.1

z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

Provider:z-ai logoz-ai
Pricing:¥250.18 in · ¥786.28 out / 1M tokens
Context:205K tokens
Text Overall
#44

Gemma 4 26B A4B

google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:262K tokens

Gemma 4 26B A4B (free)

google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:262K tokens

Gemma 4 31B

google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Provider:google logogoogle
Pricing:¥26 in · ¥101 out / 1M tokens
Context:262K tokens
Text Overall
#66

Gemma 4 31B (free)

google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3.6 Plus

qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#77

GLM 5V Turbo

z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:203K tokens
Text Overall
#95

Trinity Large Thinking

arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

Provider:arcee-ai
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#182

Grok 4.20

x-ai/grok-4.20

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:2.0M tokens

Grok 4.20 Multi-Agent

x-ai/grok-4.20-multi-agent

Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:2.0M tokens

LLM-jp-3.1 8x13B instruct4

llm-jp/llm-jp-3.1-8x13b-instruct4

LLM-jp-3.1 8x13B instruct4 is a Japanese-specialized open Mixture-of-Experts model from Japan's National Institute of Informatics (NII), with 73B total and 22B active parameters.

Provider:llm-jp
Pricing:¥16 in · ¥79 out / 1M tokens
Context:4K tokens

MiMo-V2-Pro

xiaomi/mimo-v2-pro

MiMo-V2-Pro is Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M context length, deeply optimized for agentic scenarios. It is highly adaptable to general agent frameworks like OpenClaw.

Provider:xiaomi
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#71

MiniMax M2.7

minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.

Provider:minimax logominimax
Pricing:¥53.61 in · ¥214.44 out / 1M tokens
Context:205K tokens
Text Overall
#124

GPT-5.4 Mini

openai/gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

GPT-5.4 Nano

openai/gpt-5.4-nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

GLM 5 Turbo

z-ai/glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:203K tokens

Grok 4.20 Beta

x-ai/grok-4.20-beta

Grok 4.20 Beta is xAI's newest flagship model with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truth…

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:2.0M tokens

Nemotron 3 Super

nvidia/nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications.

Provider:nvidia logonvidia
Pricing:Not on FastMetal yet
Context:262K tokens

GPT-5.4

openai/gpt-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling high-context reasoning, codi…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.1M tokens
Text Overall
#46

GPT-5.4 Pro

openai/gpt-5.4-pro

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.1M tokens

Mercury 2

inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs.…

Provider:inception
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#209

Gemini 3.1 Flash Lite Preview

google/gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#96

GPT-5.3 Chat

openai/gpt-5.3-chat

GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens

Nano Banana 2 (Gemini 3.1 Flash Image Preview)

google/gemini-3.1-flash-image-preview

Gemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:66K tokens

Gemini 3.1 Pro Preview Custom Tools

google/gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party or user-defined functions are available.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3.5-122B-A10B

qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#122

Qwen3.5-27B

qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#138

Qwen3.5-35B-A3B

qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#153

Qwen3.5-Flash

qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#150

GPT-5.3-Codex

openai/gpt-5.3-codex

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

Gemini 3.1 Pro Preview

google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#15

Claude Sonnet 4.6

anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.

Provider:anthropic logoanthropic
Pricing:¥554.4 in · ¥2,772 out / 1M tokens
Context:1.0M tokens
Text Overall
#35

Qwen3.5 397B A17B

qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#82

MiniMax M2.5

minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office wor…

Provider:minimax logominimax
Pricing:Not on FastMetal yet
Context:197K tokens
Text Overall
#159

MiniMax M2.5 (free)

minimax/minimax-m2.5:free

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office wor…

Provider:minimax logominimax
Pricing:Not on FastMetal yet
Context:197K tokens

GLM 5

z-ai/glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.

Provider:z-ai logoz-ai
Pricing:¥178.7 in · ¥571.84 out / 1M tokens
Context:205K tokens
Text Overall
#56

Claude Opus 4.6

anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refa…

Provider:anthropic logoanthropic
Pricing:¥850 in · ¥4,200 out / 1M tokens
Context:1.0M tokens
Text Overall
#6

Step 3.5 Flash

stepfun/step-3.5-flash

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.

Provider:stepfun logostepfun
Pricing:Not on FastMetal yet
Context:256K tokens
Text Overall
#156

Step 3.5 Flash (free)

stepfun/step-3.5-flash:free

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.

Provider:stepfun logostepfun
Pricing:Not on FastMetal yet
Context:256K tokens

Kimi K2.5

moonshotai/kimi-k2.5

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm.

Provider:moonshotai logomoonshotai
Pricing:Not on FastMetal yet
Context:262K tokens

Trinity Large Preview (free)

arcee-ai/trinity-large-preview:free

Trinity-Large-Preview is a frontier-scale open-weight language model from Arcee, built as a 400B-parameter sparse Mixture-of-Experts with 13B active parameters per token using 4-of-256 expert routing.

Provider:arcee-ai
Pricing:Not on FastMetal yet
Context:131K tokens

MiniMax M2-her

minimax/minimax-m2-her

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations.

Provider:minimax logominimax
Pricing:Not on FastMetal yet
Context:66K tokens

GLM 4.7 Flash

z-ai/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaborati…

Provider:z-ai logoz-ai
Pricing:¥10.722 in · ¥71.48 out / 1M tokens
Context:203K tokens
Text Overall
#184

GPT-5.2-Codex

openai/gpt-5.2-codex

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

Molmo2 8B

allenai/molmo-2-8b

Molmo2-8B is an open vision-language model developed by the Allen Institute for AI (Ai2) as part of the Molmo2 family, supporting image, video, and multi-image understanding and grounding.

Provider:allenai
Pricing:Not on FastMetal yet
Context:37K tokens
Text Overall
#233

Olmo 3.1 32B Instruct

allenai/olmo-3.1-32b-instruct

Olmo 3.1 32B Instruct is a large-scale, 32-billion-parameter instruction-tuned language model engineered for high-performance conversational AI, multi-turn dialogue, and practical instruction following.

Provider:allenai
Pricing:Not on FastMetal yet
Context:66K tokens
Text Overall
#230

MiniMax M2.1

minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development.

Provider:minimax logominimax
Pricing:Not on FastMetal yet
Context:197K tokens

GLM 4.7

z-ai/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.

Provider:z-ai logoz-ai
Pricing:¥107.22 in · ¥393.14 out / 1M tokens
Context:203K tokens
Text Overall
#81

Gemini 3 Flash Preview

google/gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Olmo 3.1 32B Think

allenai/olmo-3.1-32b-think

Olmo 3.1 32B Think is a large-scale, 32-billion-parameter model designed for deep reasoning, complex multi-step logic, and advanced instruction following.

Provider:allenai
Pricing:Not on FastMetal yet
Context:66K tokens
Text Overall
#284

MiMo-V2-Flash

xiaomi/mimo-v2-flash

MiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It is a Mixture-of-Experts model with 309B total parameters and 15B active parameters, adopting hybrid attention architecture.

Provider:xiaomi
Pricing:Not on FastMetal yet
Context:262K tokens

Nemotron 3 Nano 30B A3B

nvidia/nemotron-3-nano-30b-a3b

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems.

Provider:nvidia logonvidia
Pricing:Not on FastMetal yet
Context:262K tokens

GPT-5.2

openai/gpt-5.2

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens
Text Overall
#88

GPT-5.2 Chat

openai/gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens

GPT-5.2 Pro

openai/gpt-5.2-pro

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

Devstral 2 2512

mistralai/devstral-2512

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:262K tokens

DeepSeek V3.1 Nex N1

nex-agi/deepseek-v3.1-nex-n1

DeepSeek V3.1 Nex-N1 is the flagship release of the Nex-N1 series — a post-trained model designed to highlight agent autonomy, tool use, and real-world productivity.

Provider:nex-agi
Pricing:Not on FastMetal yet
Context:131K tokens

GLM 4.6V

z-ai/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#174

GPT-5.1-Codex-Max

openai/gpt-5.1-codex-max

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

Nova 2 Lite

amazon/nova-2-lite-v1

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text.

Provider:amazon logoamazon
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#222

DeepSeek V3.2

deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#108

DeepSeek V3.2 Speciale

deepseek/deepseek-v3.2-speciale

DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens

Mistral Large 3 2512

mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:262K tokens

INTELLECT-3

prime-intellect/intellect-3

INTELLECT-3 is a 106B-parameter Mixture-of-Experts model (12B active) post-trained from GLM-4.5-Air-Base using supervised fine-tuning (SFT) followed by large-scale reinforcement learning (RL).

Provider:prime-intellect
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#195

Claude Opus 4.5

anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#40

Olmo 3 32B Think

allenai/olmo-3-32b-think

Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios.

Provider:allenai
Pricing:Not on FastMetal yet
Context:66K tokens
Text Overall
#263

Nano Banana Pro (Gemini 3 Pro Image Preview)

google/gemini-3-pro-image-preview

Nano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthe…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:66K tokens

Grok 4.1 Fast

x-ai/grok-4.1-fast

Grok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window. Reasoning can be enabled/disabled using the `reasoning` `enabled` parameter in the API.

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:2.0M tokens

Gemini 3 Pro Preview

google/gemini-3-pro-preview

Gemini 3 Pro is Google’s flagship frontier model for high-precision multimodal reasoning, combining strong performance across text, image, video, audio, and code with a 1M-token context window.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

GPT-5.1

openai/gpt-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens
Text Overall
#84

GPT-5.1 Chat

openai/gpt-5.1-chat

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens

GPT-5.1-Codex

openai/gpt-5.1-codex

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

GPT-5.1-Codex-Mini

openai/gpt-5.1-codex-mini

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

KAT-Coder-Pro V1

kwaipilot/kat-coder-pro

KAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KAT-Coder series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving 73.4% solve rate on the SWE-Bench Ve…

Provider:kwaipilot
Pricing:Not on FastMetal yet
Context:256K tokens

Kimi K2 Thinking

moonshotai/kimi-k2-thinking

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.

Provider:moonshotai logomoonshotai
Pricing:Not on FastMetal yet
Context:131K tokens

MiniMax M2

minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,…

Provider:minimax logominimax
Pricing:Not on FastMetal yet
Context:197K tokens
Text Overall
#212

Claude Haiku 4.5

anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models.

Provider:anthropic logoanthropic
Pricing:¥184.8 in · ¥924 out / 1M tokens
Context:200K tokens
Text Overall
#131

Llama 3.3 Nemotron Super 49B V1.5

nvidia/llama-3.3-nemotron-super-49b-v1.5

Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context.

Provider:nvidia logonvidia
Pricing:Not on FastMetal yet
Context:131K tokens

Nano Banana (Gemini 2.5 Flash Image)

google/gemini-2.5-flash-image

Gemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens

GLM 4.6

z-ai/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:203K tokens
Text Overall
#110

GLM 4.6 (exacto)

z-ai/glm-4.6:exacto

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:205K tokens

Claude Sonnet 4.5

anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#62

DeepSeek V3.2 Exp

deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#114

Gemini 2.5 Flash Lite Preview 09-2025

google/gemini-2.5-flash-lite-preview-09-2025

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3 Max

qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3 VL 235B A22B Instruct

qwen/qwen3-vl-235b-a22b-instruct

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#128

Qwen3 VL 235B A22B Thinking

qwen/qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#154

DeepSeek V3.1 Terminus

deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#126

DeepSeek V3.1 Terminus (exacto)

deepseek/deepseek-v3.1-terminus:exacto

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens

Grok 4 Fast

x-ai/grok-4-fast

Grok 4 Fast is xAI's latest multimodal model with SOTA cost-efficiency and a 2M token context window. It comes in two flavors: non-reasoning and reasoning. Read more about the model on xAI's [news post](http://x.ai/news/grok-4-fast).

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:2.0M tokens

Qwen3 Next 80B A3B Instruct

qwen/qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#148

Qwen3 Next 80B A3B Instruct (free)

qwen/qwen3-next-80b-a3b-instruct:free

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3 Next 80B A3B Thinking

qwen/qwen3-next-80b-a3b-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#183

LongCat Flash Chat

meituan/longcat-flash-chat

LongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input.

Provider:meituan
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#145

Kimi K2 0905

moonshotai/kimi-k2-0905

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass.…

Provider:moonshotai logomoonshotai
Pricing:Not on FastMetal yet
Context:131K tokens

Qwen3 30B A3B Thinking 2507

qwen/qwen3-30b-a3b-thinking-2507

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens

Grok Code Fast 1

x-ai/grok-code-fast-1

Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. With reasoning traces visible in the response, developers can steer Grok Code for high-quality work flows.

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:256K tokens

DeepSeek V3.1

deepseek/deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#121

Mistral Medium 3.1

mistralai/mistral-medium-3.1

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:131K tokens

GLM 4.5V

z-ai/glm-4.5v

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understandin…

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:66K tokens
Text Overall
#199

GPT-5

openai/gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes us…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

GPT-5 Chat

openai/gpt-5-chat

GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#104

GPT-5 Mini

openai/gpt-5-mini

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

GPT-5 Nano

openai/gpt-5-nano

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:400K tokens

Claude Opus 4.1

anthropic/claude-opus-4.1

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#73

gpt-oss-120b

openai/gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.

Provider:openai logoopenai
Pricing:¥16 in · ¥79 out / 1M tokens
Context:131K tokens
Text Overall
#200

gpt-oss-120b (exacto)

openai/gpt-oss-120b:exacto

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:131K tokens

gpt-oss-120b (free)

openai/gpt-oss-120b:free

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:131K tokens

gpt-oss-20b

openai/gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deplo…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#251

gpt-oss-20b (free)

openai/gpt-oss-20b:free

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deplo…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:131K tokens

Qwen3 Coder 30B A3B Instruct

qwen/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:160K tokens

Qwen3 30B A3B Instruct 2507

qwen/qwen3-30b-a3b-instruct-2507

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#171

GLM 4.5

z-ai/glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#134

GLM 4.5 Air

z-ai/glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#180

GLM 4.5 Air (free)

z-ai/glm-4.5-air:free

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size.

Provider:z-ai logoz-ai
Pricing:Not on FastMetal yet
Context:131K tokens

Qwen3 235B A22B Thinking 2507

qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#147

Qwen3 Coder 480B A35B

qwen/qwen3-coder

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Gemini 2.5 Flash Lite

google/gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3 235B A22B Instruct 2507

qwen/qwen3-235b-a22b-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#112

Kimi K2 0711

moonshotai/kimi-k2

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass.

Provider:moonshotai logomoonshotai
Pricing:Not on FastMetal yet
Context:131K tokens

Devstral Medium

mistralai/devstral-medium

Devstral Medium is a high-performance code generation and agentic reasoning model developed jointly by Mistral AI and All Hands AI.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:131K tokens

Grok 4

x-ai/grok-4

Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs.

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:256K tokens

DeepSeek R1T2 Chimera

tngtech/deepseek-r1t2-chimera

DeepSeek-TNG-R1T2-Chimera is the second-generation Chimera model from TNG Tech. It is a 671 B-parameter mixture-of-experts text-generation model assembled from DeepSeek-AI’s R1-0528, R1, and V3-0324 checkpoints with an Assembly-of-Experts m…

Provider:tngtech
Pricing:Not on FastMetal yet
Context:164K tokens

Mercury

inception/mercury

Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their perfor…

Provider:inception
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#261

Gemini 2.5 Flash

google/gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#136

Gemini 2.5 Pro

google/gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#75

MiniMax M1

minimax/minimax-m1

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (MoE) architecture paired with a custom "lightning attention" mechanism, allowing…

Provider:minimax logominimax
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#188

Grok 3

x-ai/grok-3

Grok 3 is the latest model from xAI. It's their flagship model that excels at enterprise use cases like data extraction, coding, and text summarization. Possesses deep domain knowledge in finance, healthcare, law, and science.

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:131K tokens

Grok 3 Mini

x-ai/grok-3-mini

A lightweight model that thinks before responding. Fast, smart, and great for logic-based tasks that do not require deep domain knowledge. The raw thinking traces are accessible.

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:131K tokens

Gemini 2.5 Pro Preview 06-05

google/gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

R1 0528

deepseek/deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#115

Claude Opus 4

anthropic/claude-opus-4

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#129

Claude Sonnet 4

anthropic/claude-sonnet-4

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#160

Gemma 3n 4B

google/gemma-3n-e4b-it

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#250

Gemma 3n 4B (free)

google/gemma-3n-e4b-it:free

Gemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:8K tokens

Gemini 2.5 Pro Preview 05-06

google/gemini-2.5-pro-preview-05-06

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Mistral Medium 3

mistralai/mistral-medium-3

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:131K tokens

Mercury Coder

inception/mercury-coder

Mercury Coder is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like Claude 3.5 Haiku and GPT-4o Mini while matching their p…

Provider:inception
Pricing:Not on FastMetal yet
Context:128K tokens

Qwen3 235B A22B

qwen/qwen3-235b-a22b

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#176

Qwen3 30B A3B

qwen/qwen3-30b-a3b

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:41K tokens
Text Overall
#234

Qwen3 32B

qwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:41K tokens
Text Overall
#207

o3

openai/o3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:200K tokens

o4 Mini

openai/o4-mini

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:200K tokens

GPT-4.1

openai/gpt-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.0M tokens

GPT-4.1 Mini

openai/gpt-4.1-mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.0M tokens

GPT-4.1 Nano

openai/gpt-4.1-nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, a…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:1.0M tokens

Grok 3 Mini Beta

x-ai/grok-3-mini-beta

Grok 3 Mini is a lightweight, smaller thinking model. Unlike traditional models that generate answers immediately, Grok 3 Mini thinks before responding.

Provider:x-ai logox-ai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#193

Llama 3.1 Nemotron Ultra 253B v1

nvidia/llama-3.1-nemotron-ultra-253b-v1

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks.

Provider:nvidia logonvidia
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#206

Llama 4 Maverick

meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward pass (400B total).

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:1.0M tokens

Llama 4 Scout

meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:328K tokens

DeepSeek V3 0324

deepseek/deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#151

Qwen2.5 VL 32B Instruct

qwen/qwen2.5-vl-32b-instruct

Qwen2.5-VL-32B is a multimodal vision-language model fine-tuned through reinforcement learning for enhanced mathematical reasoning, structured outputs, and visual problem-solving capabilities.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:128K tokens

Mistral Small 3.1 24B

mistralai/mistral-small-3.1-24b-instruct

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:128K tokens

Olmo 2 32B Instruct

allenai/olmo-2-0325-32b-instruct

OLMo-2 32B Instruct is a supervised instruction-finetuned variant of the OLMo-2 32B March 2025 base model. It excels in complex reasoning and instruction-following tasks across diverse benchmarks such as GSM8K, MATH, IFEval, and general NLP…

Provider:allenai
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#305

Command A

cohere/command-a

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases.

Provider:cohere logocohere
Pricing:Not on FastMetal yet
Context:256K tokens

Gemma 3 12B

google/gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#218

Gemma 3 12B (free)

google/gemma-3-12b-it:free

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens

Gemma 3 4B

google/gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#268

Gemma 3 4B (free)

google/gemma-3-4b-it:free

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:33K tokens

Gemma 3 27B

google/gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#186

Gemma 3 27B (free)

google/gemma-3-27b-it:free

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:131K tokens

QwQ 32B

qwen/qwq-32b

QwQ is the reasoning model of the Qwen series. Compared with conventional instruction-tuned models, QwQ, which is capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problem…

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#223

Gemini 2.0 Flash Lite

google/gemini-2.0-flash-lite-001

Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5), all…

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens

Claude 3.7 Sonnet

anthropic/claude-3.7-sonnet

Claude 3.7 Sonnet is an advanced large language model with improved reasoning, coding, and problem-solving capabilities. It introduces a hybrid reasoning approach, allowing users to choose between rapid responses and extended, step-by-step…

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#181

o3 Mini High

openai/o3-mini-high

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#189

Gemini 2.0 Flash

google/gemini-2.0-flash-001

Gemini Flash 2.0 offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5).

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#191

Qwen2.5 VL 72B Instruct

qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens

Qwen-Max

qwen/qwen-max

Qwen-Max, based on Qwen2.5, provides the best inference performance among [Qwen models](/qwen), especially for complex multi-step tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens

Qwen-Plus

qwen/qwen-plus

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen VL Max

qwen/qwen-vl-max

Qwen VL Max is a visual understanding model with 7500 tokens context length. It excels in delivering optimal performance for a broader spectrum of complex tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens

o3 Mini

openai/o3-mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#205

Mistral Small 3

mistralai/mistral-small-24b-instruct-2501

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local…

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#292

R1 Distill Qwen 32B

deepseek/deepseek-r1-distill-qwen-32b

DeepSeek R1 Distill Qwen 32B is a distilled large language model based on [Qwen 2.5 32B](https://huggingface.co/Qwen/Qwen2.5-32B), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:33K tokens

R1 Distill Llama 70B

deepseek/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:131K tokens

R1

deepseek/deepseek-r1

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:64K tokens
Text Overall
#149

Phi 4

microsoft/phi-4

[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed.

Provider:microsoft logomicrosoft
Pricing:Not on FastMetal yet
Context:16K tokens
Text Overall
#304

DeepSeek V3

deepseek/deepseek-chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#192

o1

openai/o1

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:200K tokens

Command R7B (12-2024)

cohere/command-r7b-12-2024

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning and multiple steps.

Provider:cohere logocohere
Pricing:Not on FastMetal yet
Context:128K tokens

Llama 3.3 70B Instruct

meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#249

Llama 3.3 70B Instruct (free)

meta-llama/llama-3.3-70b-instruct:free

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:128K tokens

Nova Lite 1.0

amazon/nova-lite-v1

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output.

Provider:amazon logoamazon
Pricing:Not on FastMetal yet
Context:300K tokens

Nova Micro 1.0

amazon/nova-micro-v1

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost.

Provider:amazon logoamazon
Pricing:Not on FastMetal yet
Context:128K tokens

Nova Pro 1.0

amazon/nova-pro-v1

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks.

Provider:amazon logoamazon
Pricing:Not on FastMetal yet
Context:300K tokens

Mistral Large 2407

mistralai/mistral-large-2407

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#254

Mistral Large 2411

mistralai/mistral-large-2411

Mistral Large 2 2411 is an update of [Mistral Large 2](/mistralai/mistral-large) released together with [Pixtral Large 2411](/mistralai/pixtral-large-2411) It provides a significant upgrade on the previous [Mistral Large 24.07](/mistralai/m…

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#266

Pixtral Large 2411

mistralai/pixtral-large-2411

Pixtral Large is a 124B parameter, open-weight, multimodal model built on top of [Mistral Large 2](/mistralai/mistral-large-2411). The model is able to understand documents, charts and natural images.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:131K tokens

Qwen2.5 Coder 32B Instruct

qwen/qwen-2.5-coder-32b-instruct

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reaso…

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#295

Claude 3.5 Haiku

anthropic/claude-3.5-haiku

Claude 3.5 Haiku features offers enhanced capabilities in speed, coding accuracy, and tool use. Engineered to excel in real-time applications, it delivers quick response times that are essential for dynamic tasks such as chat interactions a…

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#238

Claude 3.5 Sonnet

anthropic/claude-3.5-sonnet

New Claude 3.5 Sonnet delivers better-than-Opus capabilities, faster-than-Sonnet speeds, at the same Sonnet prices. Sonnet is particularly good at: - Coding: Scores ~49% on SWE-Bench Verified, higher than the last best score, and without an…

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#179

Llama 3.1 Nemotron 70B Instruct

nvidia/llama-3.1-nemotron-70b-instruct

NVIDIA's Llama 3.1 Nemotron 70B is a language model designed for generating precise and useful responses. Leveraging [Llama 3.1 70B](/models/meta-llama/llama-3.1-70b-instruct) architecture and Reinforcement Learning from Human Feedback (RLH…

Provider:nvidia logonvidia
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#271

Llama 3.2 1B Instruct

meta-llama/llama-3.2-1b-instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:60K tokens
Text Overall
#382

Llama 3.2 3B Instruct

meta-llama/llama-3.2-3b-instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#353

Llama 3.2 3B Instruct (free)

meta-llama/llama-3.2-3b-instruct:free

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens

Qwen2.5 72B Instruct

qwen/qwen-2.5-72b-instruct

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized…

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#270

Command R (08-2024)

cohere/command-r-08-2024

command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and tool use.

Provider:cohere logocohere
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#306

Command R+ (08-2024)

cohere/command-r-plus-08-2024

command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint the same.…

Provider:cohere logocohere
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#290

GPT-4o (2024-08-06)

openai/gpt-4o-2024-08-06

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/).

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#226

Llama 3.1 405B (base)

meta-llama/llama-3.1-405b

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This is the base 405B pre-trained version. It has demonstrated strong performance compared to leading closed-source models in human evaluations.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:33K tokens

Llama 3.1 405B Instruct

meta-llama/llama-3.1-405b-instruct

The highly anticipated 400B class of Llama3 is here! Clocking in at 128k context with impressive eval scores, the Meta AI team continues to push the frontier of open-source LLMs.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens

Llama 3.1 70B Instruct

meta-llama/llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#274

Llama 3.1 8B Instruct

meta-llama/llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:16K tokens
Text Overall
#327

GPT-4o-mini

openai/gpt-4o-mini

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens

GPT-4o-mini (2024-07-18)

openai/gpt-4o-mini-2024-07-18

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#248

Gemma 2 27B

google/gemma-2-27b-it

Gemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini).

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#277

Gemma 2 9B

google/gemma-2-9b-it

Gemma 2 9B by Google is an advanced, open-source language model that sets a new standard for efficiency and performance in its size class.

Provider:google logogoogle
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#297

Mistral 7B Instruct

mistralai/mistral-7b-instruct

A high-performing, industry-standard 7.3B parameter model, with optimizations for speed and context length. *Mistral 7B Instruct has multiple version variants, and this is intended to be the latest version.*

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#383

Mistral 7B Instruct v0.3

mistralai/mistral-7b-instruct-v0.3

A high-performing, industry-standard 7.3B parameter model, with optimizations for speed and context length. An improved version of [Mistral 7B Instruct v0.2](/models/mistralai/mistral-7b-instruct-v0.2), with the following changes: - Extende…

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:33K tokens

GPT-4o

openai/gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as fast and 50% more cost-effec…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens

GPT-4o (2024-05-13)

openai/gpt-4o-2024-05-13

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as fast and 50% more cost-effec…

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#211

Llama 3 70B Instruct

meta-llama/llama-3-70b-instruct

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high quality dialogue usecases.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#289

Llama 3 8B Instruct

meta-llama/llama-3-8b-instruct

Meta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue usecases.

Provider:meta-llama logometa-llama
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#320

Mixtral 8x22B Instruct

mistralai/mixtral-8x22b-instruct

Mistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparalleled cost efficiency for its size.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:66K tokens

GPT-4 Turbo

openai/gpt-4-turbo

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens

Claude 3 Haiku

anthropic/claude-3-haiku

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance.

Provider:anthropic logoanthropic
Pricing:Not on FastMetal yet
Context:200K tokens
Text Overall
#301

Mistral Large

mistralai/mistral-large

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:128K tokens

Mixtral 8x7B Instruct

mistralai/mixtral-8x7b-instruct

Mixtral 8x7B Instruct is a pretrained generative Sparse Mixture of Experts, by Mistral AI, for chat and instruction use. Incorporates 8 experts (feed-forward networks) for a total of 47 billion parameters.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:33K tokens

GPT-4 Turbo (older v1106)

openai/gpt-4-1106-preview

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to April 2023.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#257

Mistral 7B Instruct v0.1

mistralai/mistral-7b-instruct-v0.1

A 7.3B parameter model that outperforms Llama 2 13B on all benchmarks, with optimizations for speed and context length.

Provider:mistralai logomistralai
Pricing:Not on FastMetal yet
Context:3K tokens

GPT-3.5 Turbo

openai/gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:16K tokens

GPT-4

openai/gpt-4

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning capabilities.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:8K tokens

GPT-4 (older v0314)

openai/gpt-4-0314

GPT-4-0314 is the first version of GPT-4 released, with a context length of 8,192 tokens, and was supported until June 14. Training data: up to Sep 2021.

Provider:openai logoopenai
Pricing:Not on FastMetal yet
Context:8K tokens
Text Overall
#280

deepseek-v4-flash-0731

deepseek/deepseek-v4-flash-0731

Global: DeepSeek's V4 Flash 0731

Provider:deepseek logodeepseek
Pricing:¥39.314 in · ¥117.942 out / 1M tokens
Context:- tokens

Gemini Flash Lite

google/gemini-flash-lite

Google's fastest and most cost-efficient model in the Gemini series. Delivers frontier-class performance with 2.5x faster time-to-first-token, ideal for high-volume, latency-sensitive applications.

Provider:google logogoogle
Pricing:¥0 in · ¥0 out / 1M tokens
Context:1.0M tokens

Lustify SDXL

venice/lustify-sdxl

Global: Generate animated images

Provider:venice
Pricing:Not on FastMetal yet
Context:- tokens

Nex-N2.5-Mini

nex-agi/nex-n2.5-mini

Global: Nex AGI Nex-N2.5-Mini - agentic coding model, free to try

Provider:nex-agi
Pricing:¥0 in · ¥0 out / 1M tokens
Context:262K tokens

qwen3.6-27b

qwen/qwen3.6-27b

Global: Qwen 3.6 27B

Provider:qwen logoqwen
Pricing:¥80.415 in · ¥482.49 out / 1M tokens
Context:- tokens

Qwen3.6 35B-A3B

qwen/qwen3.6-35b-a3b

Japan: Qwen 3.6 35B-A3B - hosted and run in Japan

Provider:qwen logoqwen
Pricing:¥32 in · ¥158 out / 1M tokens
Context:- tokens

Qwen3 Coder 480B A35B

qwen/qwen3-coder-480b-a35b-instruct-fp8

Alibaba's most capable open-source agentic coding model. A Mixture-of-Experts architecture with 480B total parameters (35B active), trained on 7.5 trillion tokens.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

random-free

openrouter/free

Global: Random - free to try

Provider:openrouter logoopenrouter
Pricing:¥0 in · ¥0 out / 1M tokens
Context:200K tokens

Seedream 4.5

bytedance-seed/seedream-4.5

Global: Generate images

Provider:bytedance-seed logobytedance-seed
Pricing:¥7.15 per image
Context:- tokens

Voxtral Mini 3B

mistralai/voxtral-mini-3b-2507

A 3B-parameter speech-language model built on the Ministral-3B backbone with an audio encoder for state-of-the-art audio understanding. Supports speech transcription, translation, audio Q&A, and voice-to-function calling across 8 languages.

Provider:mistralai logomistralai
Pricing:¥8.4 in · ¥8.4 out / 1M tokens
Context:33K tokens

Z-Image Turbo

tongyi/z-image-turbo

Global: Generate realistic images

Provider:tongyi
Pricing:¥1.68 per image
Context:- tokens