303 models
DeepSeek V4.1 Flash
deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
Mercury 2.5
inception/mercury-2.5Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury 2.5 produces and refines multiple tokens in parallel, achieving...
GPT-6 Astra
openai/gpt-6-astraGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strengths in long-horizon...
GPT-6 Astra Pro
openai/gpt-6-astra-proGPT-6 Astra Pro is the same underlying model as [GPT-6 Astra](https://openrouter.ai/openai/gpt-6-astra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks.
Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...
Gemini 3.8 Flash
google/gemini-3.8-flashGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning.
Muse Spark 1.3
meta/muse-spark-1.3Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track of information across extended tasks, work through...
Muse Spark 1.3 Contributor
meta/muse-spark-1.3-contributorMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and early-stage agentic, multi-agent, and coding workflows. It is designed to track information...
Claude Fable 5.1
anthropic/claude-fable-5.1Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and knowledge work: long code refactors, front-end and visual...
Hy4 preview
tencent/hy4-previewTencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents, complex tool-use workflows, and productivity tasks that...
GLM 5.3 Flash
z-ai/glm-5.3-flashGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear attention architecture maintains accurate long-context behavior while...
Qwen3.8 Flash
qwen/qwen3.8-flashQwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.
Muse Spark 1.2 Contributor
meta/muse-spark-1.2-contributorMuse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s meaningfully cheaper than Muse Spark...
GLM 5.3
z-ai/glm-5.3GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and output with a 1M-token context window, and improves...
Qwen3.8 27B
qwen/qwen3.8-27bQwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...
Gemini 3.7 Flash
google/gemini-3.7-flashGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks that require responsive performance and reliable multi-step...
Grok 4.6
x-ai/grok-4.6Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Qwen3.8 2.4T A95B
qwen/qwen3.8-2.4t-a95bQwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...
Muse Glimmer 30B
meta/muse-glimmer-30bMeta's dense, open-weight 30B multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for long-horizon autonomous-agent tasks on consumer hardware.
Solar Pro 4
upstage/solar-pro4Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic workflows, with strong capabilities in office productivity, document-intensive...
Muse Spark 1.2
meta/muse-spark-1.2Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, audio, and PDF documents, returns text, and offers a 1M-token context...
Qwen3.8 Max
qwen/qwen3.8-maxAlibaba's flagship Qwen3.8 model: a 2.4T-parameter MoE (95B active) multimodal reasoning model with a 1M-token context window.
Claude Opus 5
anthropic/claude-opus-5Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end software tasks, code review and bug finding, visual analysis...
Gemini 3.5 Flash Lite
google/gemini-3.5-flash-liteGemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
Gemini 3.6 Flash
google/gemini-3.6-flashGemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...
Inkling
thinkingmachines/inklingInkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is designed for general-purpose reasoning, coding, agentic and tool-use systems,...
Kimi K3
moonshotai/kimi-k3Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Muse Spark 1.1
meta/muse-spark-1.1Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, audio, and PDF documents and returns text, with a 1M-token context...
GPT-5.6 Luna
openai/gpt-5.6-lunaGPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat, classification, and lightweight agentic workflows, providing capable reasoning for...
GPT-5.6 Sol
openai/gpt-5.6-solGPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is particularly strong at command-line and multi-step coding tasks...
GPT-5.6 Terra
openai/gpt-5.6-terraGPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is suited for everyday coding, reasoning, and agentic...
Grok 4.5
x-ai/grok-4.5Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
Hy3
tencent/hy3Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
Laguna XS 2.1
poolside/laguna-xs-2.1Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Laguna XS 2.1 (free)
poolside/laguna-xs-2.1:freeLaguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
Claude Sonnet 5
anthropic/claude-sonnet-5Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...
Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)
google/gemini-3.1-flash-lite-imageNano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's fastest, most cost-efficient Gemini image model, built for high-velocity developer pipelines and rapid-fire visual exploration. It delivers text-to-image generation...
Nano Banana 2 (Gemini 3.1 Flash Image)
google/gemini-3.1-flash-imageGoogle's fast, high-quality image generation model.
Nano Banana Pro (Gemini 3 Pro Image)
google/gemini-3-pro-imageNano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and...
GLM 5.2
z-ai/glm-5.2GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
Fusion
openrouter/fusionFusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...
Kimi K2.7 Code
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over long contexts. It uses a native multimodal mixture-of-experts...
Claude Fable 5
anthropic/claude-fable-5Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...
Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55bNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
Qwen3.7 Plus
qwen/qwen3.7-plusQwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
MiniMax M3
minimax/minimax-m3MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
Step 3.7 Flash
stepfun/step-3.7-flashStep 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision encoder for native image and video understanding, activating roughly 11B parameters...
Claude Opus 4.8
anthropic/claude-opus-4.8Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
Claude Opus 4.8 (Fast)
anthropic/claude-opus-4.8-fastFast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8.
Qwen3.7 Max
qwen/qwen3.7-maxQwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...
Gemini 3.5 Flash
google/gemini-3.5-flashGemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...
Claude Opus 4.7 (Fast)
anthropic/claude-opus-4.7-fastFast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
Gemini 3.1 Flash Lite
google/gemini-3.1-flash-liteGemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...
Granite 4.1 8B
ibm-granite/granite-4.1-8bGranite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window and is designed for enterprise tasks...
Grok 4.3
x-ai/grok-4.3Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
Mistral Medium 3.5
mistralai/mistral-medium-3-5Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
Laguna M.1
poolside/laguna-m.1Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...
Laguna M.1 (free)
poolside/laguna-m.1:freeLaguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...
Google Gemini Pro Latest
~google/gemini-pro-latestThis model always redirects to the latest model in the Google Gemini Pro family.
Qwen3.6 Max Preview
qwen/qwen3.6-max-previewQwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...
DeepSeek V4 Flash
deepseek/deepseek-v4-flashDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
DeepSeek V4 Pro
deepseek/deepseek-v4-proDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
GPT-5.5
openai/gpt-5.5GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...
GPT-5.5 Pro
openai/gpt-5.5-proGPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token context window (922K input, 128K output) with support for...
Hy3 preview
tencent/hy3-previewHy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable reasoning levels across disabled, low, and high modes, allowing it to...
MiMo-V2.5
xiaomi/mimo-v2.5MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing MiMo-V2-Omni in multimodal perception across image and video understanding...
MiMo-V2.5-Pro
xiaomi/mimo-v2.5-proMiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and long-horizon tasks, with top rankings on benchmarks such as ClawEval, GDPVal, and SWE-bench Pro....
GPT-5.4 Image 2
openai/gpt-5.4-image-2[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2.
Kimi K2.6
moonshotai/kimi-k2.6Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent orchestration. It handles complex end-to-end coding tasks across Python, Rust, and Go, and...
Claude Opus 4.7
anthropic/claude-opus-4.7Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...
GLM 5.1
z-ai/glm-5.1GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...
Gemma 4 26B A4B
google/gemma-4-26b-a4b-itGemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemma 4 26B A4B (free)
google/gemma-4-26b-a4b-it:freeGemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
Gemma 4 31B
google/gemma-4-31b-itGemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Gemma 4 31B (free)
google/gemma-4-31b-it:freeGemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
Qwen3.6 Plus
qwen/qwen3.6-plusQwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...
GLM 5V Turbo
z-ai/glm-5v-turboGLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...
Trinity Large Thinking
arcee-ai/trinity-large-thinkingTrinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
Grok 4.20
x-ai/grok-4.20Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...
Grok 4.20 Multi-Agent
x-ai/grok-4.20-multi-agentGrok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...
LLM-jp-3.1 8x13B instruct4
llm-jp/llm-jp-3.1-8x13b-instruct4LLM-jp-3.1 8x13B instruct4 is a Japanese-specialized open Mixture-of-Experts model from Japan's National Institute of Informatics (NII), with 73B total and 22B active parameters.
MiMo-V2-Pro
xiaomi/mimo-v2-proMiMo-V2-Pro is Xiaomi's flagship foundation model, featuring over 1T total parameters and a 1M context length, deeply optimized for agentic scenarios. It is highly adaptable to general agent frameworks like OpenClaw.
MiniMax M2.7
minimax/minimax-m2.7MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement.
GPT-5.4 Mini
openai/gpt-5.4-miniGPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads.
GPT-5.4 Nano
openai/gpt-5.4-nanoGPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks.
GLM 5 Turbo
z-ai/glm-5-turboGLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios.
Grok 4.20 Beta
x-ai/grok-4.20-betaGrok 4.20 Beta is xAI's newest flagship model with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering consistently precise and truth…
Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12bNVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications.
GPT-5.4
openai/gpt-5.4GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for text and image inputs, enabling high-context reasoning, codi…
GPT-5.4 Pro
openai/gpt-5.4-proGPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex, high-stakes tasks.
Mercury 2
inception/mercury-2Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs.…
Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-previewGemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on overall quality and approaches Gemini 2.5 Flash performance across key capabilities.
GPT-5.3 Chat
openai/gpt-5.3-chatGPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful.
Nano Banana 2 (Gemini 3.1 Flash Image Preview)
google/gemini-3.1-flash-image-previewGemini 3.1 Flash Image Preview, a.k.a. "Nano Banana 2," is Google’s latest state of the art image generation and editing model, delivering Pro-level visual quality at Flash speed.
Gemini 3.1 Pro Preview Custom Tools
google/gemini-3.1-pro-preview-customtoolsGemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash tool when more efficient third-party or user-defined functions are available.
Qwen3.5-122B-A10B
qwen/qwen3.5-122b-a10bThe Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
Qwen3.5-27B
qwen/qwen3.5-27bThe Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance.
Qwen3.5-35B-A3B
qwen/qwen3.5-35b-a3bThe Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency.
Qwen3.5-Flash
qwen/qwen3.5-flash-02-23The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
GPT-5.3-Codex
openai/gpt-5.3-codexGPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge capabilities of GPT-5.2.
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-previewGemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows.
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work.
Qwen3.5 397B A17B
qwen/qwen3.5-397b-a17bThe Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
MiniMax M2.5
minimax/minimax-m2.5MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office wor…
MiniMax M2.5 (free)
minimax/minimax-m2.5:freeMiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office wor…
GLM 5
z-ai/glm-5GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows.
Claude Opus 4.6
anthropic/claude-opus-4.6Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective for large codebases, complex refa…
Step 3.5 Flash
stepfun/step-3.5-flashStep 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.
Step 3.5 Flash (free)
stepfun/step-3.5-flash:freeStep 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively activates only 11B of its 196B parameters per token.
Kimi K2.5
moonshotai/kimi-k2.5Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm.
Trinity Large Preview (free)
arcee-ai/trinity-large-preview:freeTrinity-Large-Preview is a frontier-scale open-weight language model from Arcee, built as a 400B-parameter sparse Mixture-of-Experts with 13B active parameters per token using 4-of-256 expert routing.
MiniMax M2-her
minimax/minimax-m2-herMiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations.
GLM 4.7 Flash
z-ai/glm-4.7-flashAs a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning, and tool collaborati…
GPT-5.2-Codex
openai/gpt-5.2-codexGPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.
Molmo2 8B
allenai/molmo-2-8bMolmo2-8B is an open vision-language model developed by the Allen Institute for AI (Ai2) as part of the Molmo2 family, supporting image, video, and multi-image understanding and grounding.
Olmo 3.1 32B Instruct
allenai/olmo-3.1-32b-instructOlmo 3.1 32B Instruct is a large-scale, 32-billion-parameter instruction-tuned language model engineered for high-performance conversational AI, multi-turn dialogue, and practical instruction following.
MiniMax M2.1
minimax/minimax-m2.1MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development.
GLM 4.7
z-ai/glm-4.7GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution.
Gemini 3 Flash Preview
google/gemini-3-flash-previewGemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance.
Olmo 3.1 32B Think
allenai/olmo-3.1-32b-thinkOlmo 3.1 32B Think is a large-scale, 32-billion-parameter model designed for deep reasoning, complex multi-step logic, and advanced instruction following.
MiMo-V2-Flash
xiaomi/mimo-v2-flashMiMo-V2-Flash is an open-source foundation language model developed by Xiaomi. It is a Mixture-of-Experts model with 309B total parameters and 15B active parameters, adopting hybrid attention architecture.
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3bNVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems.
GPT-5.2
openai/gpt-5.2GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1.
GPT-5.2 Chat
openai/gpt-5.2-chatGPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence.
GPT-5.2 Pro
openai/gpt-5.2-proGPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro.
Devstral 2 2512
mistralai/devstral-2512Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window.
DeepSeek V3.1 Nex N1
nex-agi/deepseek-v3.1-nex-n1DeepSeek V3.1 Nex-N1 is the flagship release of the Nex-N1 series — a post-trained model designed to highlight agent autonomy, tool use, and real-world productivity.
GLM 4.6V
z-ai/glm-4.6vGLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media.
GPT-5.1-Codex-Max
openai/gpt-5.1-codex-maxGPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks.
Nova 2 Lite
amazon/nova-2-lite-v1Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text.
DeepSeek V3.2
deepseek/deepseek-v3.2DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance.
DeepSeek V3.2 Speciale
deepseek/deepseek-v3.2-specialeDeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance.
Mistral Large 3 2512
mistralai/mistral-large-2512Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
INTELLECT-3
prime-intellect/intellect-3INTELLECT-3 is a 106B-parameter Mixture-of-Experts model (12B active) post-trained from GLM-4.5-Air-Base using supervised fine-tuning (SFT) followed by large-scale reinforcement learning (RL).
Claude Opus 4.5
anthropic/claude-opus-4.5Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use.
Olmo 3 32B Think
allenai/olmo-3-32b-thinkOlmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios.
Nano Banana Pro (Gemini 3 Pro Image Preview)
google/gemini-3-pro-image-previewNano Banana Pro is Google’s most advanced image-generation and editing model, built on Gemini 3 Pro. It extends the original Nano Banana with significantly improved multimodal reasoning, real-world grounding, and high-fidelity visual synthe…
Grok 4.1 Fast
x-ai/grok-4.1-fastGrok 4.1 Fast is xAI's best agentic tool calling model that shines in real-world use cases like customer support and deep research. 2M context window. Reasoning can be enabled/disabled using the `reasoning` `enabled` parameter in the API.
Gemini 3 Pro Preview
google/gemini-3-pro-previewGemini 3 Pro is Google’s flagship frontier model for high-precision multimodal reasoning, combining strong performance across text, image, video, audio, and code with a 1M-token context window.
GPT-5.1
openai/gpt-5.1GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5.
GPT-5.1 Chat
openai/gpt-5.1-chatGPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence.
GPT-5.1-Codex
openai/gpt-5.1-codexGPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive development sessions and long, independent execution of complex engineering tasks.
GPT-5.1-Codex-Mini
openai/gpt-5.1-codex-miniGPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
KAT-Coder-Pro V1
kwaipilot/kat-coder-proKAT-Coder-Pro V1 is KwaiKAT's most advanced agentic coding model in the KAT-Coder series. Designed specifically for agentic coding tasks, it excels in real-world software engineering scenarios, achieving 73.4% solve rate on the SWE-Bench Ve…
Kimi K2 Thinking
moonshotai/kimi-k2-thinkingKimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning.
MiniMax M2
minimax/minimax-m2MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,…
Claude Haiku 4.5
anthropic/claude-haiku-4.5Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models.
Llama 3.3 Nemotron Super 49B V1.5
nvidia/llama-3.3-nemotron-super-49b-v1.5Llama-3.3-Nemotron-Super-49B-v1.5 is a 49B-parameter, English-centric reasoning/chat model derived from Meta’s Llama-3.3-70B-Instruct with a 128K context.
Nano Banana (Gemini 2.5 Flash Image)
google/gemini-2.5-flash-imageGemini 2.5 Flash Image, a.k.a. "Nano Banana," is now generally available. It is a state of the art image generation model with contextual understanding. It is capable of image generation, edits, and multi-turn conversations.
GLM 4.6
z-ai/glm-4.6Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.
GLM 4.6 (exacto)
z-ai/glm-4.6:exactoCompared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex agentic tasks.
Claude Sonnet 4.5
anthropic/claude-sonnet-4.5Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows.
DeepSeek V3.2 Exp
deepseek/deepseek-v3.2-expDeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures.
Gemini 2.5 Flash Lite Preview 09-2025
google/gemini-2.5-flash-lite-preview-09-2025Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.
Qwen3 Max
qwen/qwen3-maxQwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version.
Qwen3 VL 235B A22B Instruct
qwen/qwen3-vl-235b-a22b-instructQwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video.
Qwen3 VL 235B A22B Thinking
qwen/qwen3-vl-235b-a22b-thinkingQwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math.
DeepSeek V3.1 Terminus
deepseek/deepseek-v3.1-terminusDeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…
DeepSeek V3.1 Terminus (exacto)
deepseek/deepseek-v3.1-terminus:exactoDeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…
Grok 4 Fast
x-ai/grok-4-fastGrok 4 Fast is xAI's latest multimodal model with SOTA cost-efficiency and a 2M token context window. It comes in two flavors: non-reasoning and reasoning. Read more about the model on xAI's [news post](http://x.ai/news/grok-4-fast).
Qwen3 Next 80B A3B Instruct
qwen/qwen3-next-80b-a3b-instructQwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.
Qwen3 Next 80B A3B Instruct (free)
qwen/qwen3-next-80b-a3b-instruct:freeQwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.
Qwen3 Next 80B A3B Thinking
qwen/qwen3-next-80b-a3b-thinkingQwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default.
LongCat Flash Chat
meituan/longcat-flash-chatLongCat-Flash-Chat is a large-scale Mixture-of-Experts (MoE) model with 560B total parameters, of which 18.6B–31.3B (≈27B on average) are dynamically activated per input.
Kimi K2 0905
moonshotai/kimi-k2-0905Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass.…
Qwen3 30B A3B Thinking 2507
qwen/qwen3-30b-a3b-thinking-2507Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking.
Grok Code Fast 1
x-ai/grok-code-fast-1Grok Code Fast 1 is a speedy and economical reasoning model that excels at agentic coding. With reasoning traces visible in the response, developers can steer Grok Code for high-quality work flows.
DeepSeek V3.1
deepseek/deepseek-chat-v3.1DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates.
Mistral Medium 3.1
mistralai/mistral-medium-3.1Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost.
GLM 4.5V
z-ai/glm-4.5vGLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B parameters and 12B activated parameters, it achieves state-of-the-art results in video understandin…
GPT-5
openai/gpt-5GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy in high-stakes us…
GPT-5 Chat
openai/gpt-5-chatGPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.
GPT-5 Mini
openai/gpt-5-miniGPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost.
GPT-5 Nano
openai/gpt-5-nanoGPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments.
Claude Opus 4.1
anthropic/claude-opus-4.1Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks.
gpt-oss-120b
openai/gpt-oss-120bgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.
gpt-oss-120b (exacto)
openai/gpt-oss-120b:exactogpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.
gpt-oss-120b (free)
openai/gpt-oss-120b:freegpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases.
gpt-oss-20b
openai/gpt-oss-20bgpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deplo…
gpt-oss-20b (free)
openai/gpt-oss-20b:freegpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for lower-latency inference and deplo…
Qwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instructQwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use.
Qwen3 30B A3B Instruct 2507
qwen/qwen3-30b-a3b-instruct-2507Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference.
GLM 4.5
z-ai/glm-4.5GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens.
GLM 4.5 Air
z-ai/glm-4.5-airGLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size.
GLM 4.5 Air (free)
z-ai/glm-4.5-air:freeGLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter size.
Qwen3 235B A22B Thinking 2507
qwen/qwen3-235b-a22b-thinking-2507Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks.
Qwen3 Coder 480B A35B
qwen/qwen3-coderQwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories.
Gemini 2.5 Flash Lite
google/gemini-2.5-flash-liteGemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency.
Qwen3 235B A22B Instruct 2507
qwen/qwen3-235b-a22b-2507Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass.
Kimi K2 0711
moonshotai/kimi-k2Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass.
Devstral Medium
mistralai/devstral-mediumDevstral Medium is a high-performance code generation and agentic reasoning model developed jointly by Mistral AI and All Hands AI.
Grok 4
x-ai/grok-4Grok 4 is xAI's latest reasoning model with a 256k context window. It supports parallel tool calling, structured outputs, and both image and text inputs.
DeepSeek R1T2 Chimera
tngtech/deepseek-r1t2-chimeraDeepSeek-TNG-R1T2-Chimera is the second-generation Chimera model from TNG Tech. It is a 671 B-parameter mixture-of-experts text-generation model assembled from DeepSeek-AI’s R1-0528, R1, and V3-0324 checkpoints with an Assembly-of-Experts m…
Mercury
inception/mercuryMercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their perfor…
Gemini 2.5 Flash
google/gemini-2.5-flashGemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks.
Gemini 2.5 Pro
google/gemini-2.5-proGemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.
MiniMax M1
minimax/minimax-m1MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (MoE) architecture paired with a custom "lightning attention" mechanism, allowing…
Grok 3
x-ai/grok-3Grok 3 is the latest model from xAI. It's their flagship model that excels at enterprise use cases like data extraction, coding, and text summarization. Possesses deep domain knowledge in finance, healthcare, law, and science.
Grok 3 Mini
x-ai/grok-3-miniA lightweight model that thinks before responding. Fast, smart, and great for logic-based tasks that do not require deep domain knowledge. The raw thinking traces are accessible.
Gemini 2.5 Pro Preview 06-05
google/gemini-2.5-pro-previewGemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.
R1 0528
deepseek/deepseek-r1-0528May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens.
Claude Opus 4
anthropic/claude-opus-4Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows.
Claude Sonnet 4
anthropic/claude-sonnet-4Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability.
Gemma 3n 4B
google/gemma-3n-e4b-itGemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets.
Gemma 3n 4B (free)
google/gemma-3n-e4b-it:freeGemma 3n E4B-it is optimized for efficient execution on mobile and low-resource devices, such as phones, laptops, and tablets.
Gemini 2.5 Pro Preview 05-06
google/gemini-2.5-pro-preview-05-06Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks.
Mistral Medium 3
mistralai/mistral-medium-3Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost.
Mercury Coder
inception/mercury-coderMercury Coder is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like Claude 3.5 Haiku and GPT-4o Mini while matching their p…
Qwen3 235B A22B
qwen/qwen3-235b-a22bQwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass.
Qwen3 30B A3B
qwen/qwen3-30b-a3bQwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks.
Qwen3 32B
qwen/qwen3-32bQwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue.
o3
openai/o3o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also excels at technical writing and instruction-following.
o4 Mini
openai/o4-miniOpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and agentic capabilities.
GPT-4.1
openai/gpt-4.1GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context reasoning.
GPT-4.1 Mini
openai/gpt-4.1-miniGPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost.
GPT-4.1 Nano
openai/gpt-4.1-nanoFor tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a small size with its 1 million token context window, and scores 80.1% on MMLU, 50.3% on GPQA, a…
Grok 3 Mini Beta
x-ai/grok-3-mini-betaGrok 3 Mini is a lightweight, smaller thinking model. Unlike traditional models that generate answers immediately, Grok 3 Mini thinks before responding.
Llama 3.1 Nemotron Ultra 253B v1
nvidia/llama-3.1-nemotron-ultra-253b-v1Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks.
Llama 4 Maverick
meta-llama/llama-4-maverickLlama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward pass (400B total).
Llama 4 Scout
meta-llama/llama-4-scoutLlama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B.
DeepSeek V3 0324
deepseek/deepseek-chat-v3-0324DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.
Qwen2.5 VL 32B Instruct
qwen/qwen2.5-vl-32b-instructQwen2.5-VL-32B is a multimodal vision-language model fine-tuned through reinforcement learning for enhanced mathematical reasoning, structured outputs, and visual problem-solving capabilities.
Mistral Small 3.1 24B
mistralai/mistral-small-3.1-24b-instructMistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities.
Olmo 2 32B Instruct
allenai/olmo-2-0325-32b-instructOLMo-2 32B Instruct is a supervised instruction-finetuned variant of the OLMo-2 32B March 2025 base model. It excels in complex reasoning and instruction-following tasks across diverse benchmarks such as GSM8K, MATH, IFEval, and general NLP…
Command A
cohere/command-aCommand A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases.
Gemma 3 12B
google/gemma-3-12b-itGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…
Gemma 3 12B (free)
google/gemma-3-12b-it:freeGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…
Gemma 3 4B
google/gemma-3-4b-itGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…
Gemma 3 4B (free)
google/gemma-3-4b-it:freeGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…
Gemma 3 27B
google/gemma-3-27b-itGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…
Gemma 3 27B (free)
google/gemma-3-27b-it:freeGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities, including structu…
QwQ 32B
qwen/qwq-32bQwQ is the reasoning model of the Qwen series. Compared with conventional instruction-tuned models, QwQ, which is capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problem…
Gemini 2.0 Flash Lite
google/gemini-2.0-flash-lite-001Gemini 2.0 Flash Lite offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5), all…
Claude 3.7 Sonnet
anthropic/claude-3.7-sonnetClaude 3.7 Sonnet is an advanced large language model with improved reasoning, coding, and problem-solving capabilities. It introduces a hybrid reasoning approach, allowing users to choose between rapid responses and extended, step-by-step…
o3 Mini High
openai/o3-mini-highOpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding…
Gemini 2.0 Flash
google/gemini-2.0-flash-001Gemini Flash 2.0 offers a significantly faster time to first token (TTFT) compared to [Gemini Flash 1.5](/google/gemini-flash-1.5), while maintaining quality on par with larger models like [Gemini Pro 1.5](/google/gemini-pro-1.5).
Qwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instructQwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.
Qwen-Max
qwen/qwen-maxQwen-Max, based on Qwen2.5, provides the best inference performance among [Qwen models](/qwen), especially for complex multi-step tasks.
Qwen-Plus
qwen/qwen-plusQwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.
Qwen VL Max
qwen/qwen-vl-maxQwen VL Max is a visual understanding model with 7500 tokens context length. It excels in delivering optimal performance for a broader spectrum of complex tasks.
o3 Mini
openai/o3-miniOpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding.
Mistral Small 3
mistralai/mistral-small-24b-instruct-2501Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0 license, it features both pre-trained and instruction-tuned versions designed for efficient local…
R1 Distill Qwen 32B
deepseek/deepseek-r1-distill-qwen-32bDeepSeek R1 Distill Qwen 32B is a distilled large language model based on [Qwen 2.5 32B](https://huggingface.co/Qwen/Qwen2.5-32B), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).
R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70bDeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).
R1
deepseek/deepseek-r1DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.
Phi 4
microsoft/phi-4[Microsoft Research](/microsoft) Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or where quick responses are needed.
DeepSeek V3
deepseek/deepseek-chatDeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions.
o1
openai/o1The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained with large-scale reinforcement learning to reason using chain of thought.
Command R7B (12-2024)
cohere/command-r7b-12-2024Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar tasks requiring complex reasoning and multiple steps.
Llama 3.3 70B Instruct
meta-llama/llama-3.3-70b-instructThe Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
Llama 3.3 70B Instruct (free)
meta-llama/llama-3.3-70b-instruct:freeThe Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
Nova Lite 1.0
amazon/nova-lite-v1Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output.
Nova Micro 1.0
amazon/nova-micro-v1Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost.
Nova Pro 1.0
amazon/nova-pro-v1Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks.
Mistral Large 2407
mistralai/mistral-large-2407This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more.
Mistral Large 2411
mistralai/mistral-large-2411Mistral Large 2 2411 is an update of [Mistral Large 2](/mistralai/mistral-large) released together with [Pixtral Large 2411](/mistralai/pixtral-large-2411) It provides a significant upgrade on the previous [Mistral Large 24.07](/mistralai/m…
Pixtral Large 2411
mistralai/pixtral-large-2411Pixtral Large is a 124B parameter, open-weight, multimodal model built on top of [Mistral Large 2](/mistralai/mistral-large-2411). The model is able to understand documents, charts and natural images.
Qwen2.5 Coder 32B Instruct
qwen/qwen-2.5-coder-32b-instructQwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reaso…
Claude 3.5 Haiku
anthropic/claude-3.5-haikuClaude 3.5 Haiku features offers enhanced capabilities in speed, coding accuracy, and tool use. Engineered to excel in real-time applications, it delivers quick response times that are essential for dynamic tasks such as chat interactions a…
Claude 3.5 Sonnet
anthropic/claude-3.5-sonnetNew Claude 3.5 Sonnet delivers better-than-Opus capabilities, faster-than-Sonnet speeds, at the same Sonnet prices. Sonnet is particularly good at: - Coding: Scores ~49% on SWE-Bench Verified, higher than the last best score, and without an…
Llama 3.1 Nemotron 70B Instruct
nvidia/llama-3.1-nemotron-70b-instructNVIDIA's Llama 3.1 Nemotron 70B is a language model designed for generating precise and useful responses. Leveraging [Llama 3.1 70B](/models/meta-llama/llama-3.1-70b-instruct) architecture and Reinforcement Learning from Human Feedback (RLH…
Llama 3.2 1B Instruct
meta-llama/llama-3.2-1b-instructLlama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and multilingual text analysis.
Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instructLlama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.
Llama 3.2 3B Instruct (free)
meta-llama/llama-3.2-3b-instruct:freeLlama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization.
Qwen2.5 72B Instruct
qwen/qwen-2.5-72b-instructQwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized…
Command R (08-2024)
cohere/command-r-08-2024command-r-08-2024 is an update of the [Command R](/models/cohere/command-r) with improved performance for multilingual retrieval-augmented generation (RAG) and tool use.
Command R+ (08-2024)
cohere/command-r-plus-08-2024command-r-plus-08-2024 is an update of the [Command R+](/models/cohere/command-r-plus) with roughly 50% higher throughput and 25% lower latencies as compared to the previous Command R+ version, while keeping the hardware footprint the same.…
GPT-4o (2024-08-06)
openai/gpt-4o-2024-08-06The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the respone_format. Read more [here](https://openai.com/index/introducing-structured-outputs-in-the-api/).
Llama 3.1 405B (base)
meta-llama/llama-3.1-405bMeta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This is the base 405B pre-trained version. It has demonstrated strong performance compared to leading closed-source models in human evaluations.
Llama 3.1 405B Instruct
meta-llama/llama-3.1-405b-instructThe highly anticipated 400B class of Llama3 is here! Clocking in at 128k context with impressive eval scores, the Meta AI team continues to push the frontier of open-source LLMs.
Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instructMeta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases.
Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instructMeta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient.
GPT-4o-mini
openai/gpt-4o-miniGPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.
GPT-4o-mini (2024-07-18)
openai/gpt-4o-mini-2024-07-18GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs.
Gemma 2 27B
google/gemma-2-27b-itGemma 2 27B by Google is an open model built from the same research and technology used to create the [Gemini models](/models?q=gemini).
Gemma 2 9B
google/gemma-2-9b-itGemma 2 9B by Google is an advanced, open-source language model that sets a new standard for efficiency and performance in its size class.
Mistral 7B Instruct
mistralai/mistral-7b-instructA high-performing, industry-standard 7.3B parameter model, with optimizations for speed and context length. *Mistral 7B Instruct has multiple version variants, and this is intended to be the latest version.*
Mistral 7B Instruct v0.3
mistralai/mistral-7b-instruct-v0.3A high-performing, industry-standard 7.3B parameter model, with optimizations for speed and context length. An improved version of [Mistral 7B Instruct v0.2](/models/mistralai/mistral-7b-instruct-v0.2), with the following changes: - Extende…
GPT-4o
openai/gpt-4oGPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as fast and 50% more cost-effec…
GPT-4o (2024-05-13)
openai/gpt-4o-2024-05-13GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of [GPT-4 Turbo](/models/openai/gpt-4-turbo) while being twice as fast and 50% more cost-effec…
Llama 3 70B Instruct
meta-llama/llama-3-70b-instructMeta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 70B instruct-tuned version was optimized for high quality dialogue usecases.
Llama 3 8B Instruct
meta-llama/llama-3-8b-instructMeta's latest class of model (Llama 3) launched with a variety of sizes & flavors. This 8B instruct-tuned version was optimized for high quality dialogue usecases.
Mixtral 8x22B Instruct
mistralai/mixtral-8x22b-instructMistral's official instruct fine-tuned version of [Mixtral 8x22B](/models/mistralai/mixtral-8x22b). It uses 39B active parameters out of 141B, offering unparalleled cost efficiency for its size.
GPT-4 Turbo
openai/gpt-4-turboThe latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December 2023.
Claude 3 Haiku
anthropic/claude-3-haikuClaude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance.
Mistral Large
mistralai/mistral-largeThis is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more.
Mixtral 8x7B Instruct
mistralai/mixtral-8x7b-instructMixtral 8x7B Instruct is a pretrained generative Sparse Mixture of Experts, by Mistral AI, for chat and instruction use. Incorporates 8 experts (feed-forward networks) for a total of 47 billion parameters.
GPT-4 Turbo (older v1106)
openai/gpt-4-1106-previewThe latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to April 2023.
Mistral 7B Instruct v0.1
mistralai/mistral-7b-instruct-v0.1A 7.3B parameter model that outperforms Llama 2 13B on all benchmarks, with optimizations for speed and context length.
GPT-3.5 Turbo
openai/gpt-3.5-turboGPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional completion tasks. Training data up to Sep 2021.
GPT-4
openai/gpt-4OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous models due to its broader general knowledge and advanced reasoning capabilities.
GPT-4 (older v0314)
openai/gpt-4-0314GPT-4-0314 is the first version of GPT-4 released, with a context length of 8,192 tokens, and was supported until June 14. Training data: up to Sep 2021.
deepseek-v4-flash-0731
deepseek/deepseek-v4-flash-0731Global: DeepSeek's V4 Flash 0731
Gemini Flash Lite
google/gemini-flash-liteGoogle's fastest and most cost-efficient model in the Gemini series. Delivers frontier-class performance with 2.5x faster time-to-first-token, ideal for high-volume, latency-sensitive applications.
Lustify SDXL
venice/lustify-sdxlGlobal: Generate animated images
Nex-N2.5-Mini
nex-agi/nex-n2.5-miniGlobal: Nex AGI Nex-N2.5-Mini - agentic coding model, free to try
qwen3.6-27b
qwen/qwen3.6-27bGlobal: Qwen 3.6 27B
Qwen3.6 35B-A3B
qwen/qwen3.6-35b-a3bJapan: Qwen 3.6 35B-A3B - hosted and run in Japan
Qwen3 Coder 480B A35B
qwen/qwen3-coder-480b-a35b-instruct-fp8Alibaba's most capable open-source agentic coding model. A Mixture-of-Experts architecture with 480B total parameters (35B active), trained on 7.5 trillion tokens.
random-free
openrouter/freeGlobal: Random - free to try
Seedream 4.5
bytedance-seed/seedream-4.5Global: Generate images
Voxtral Mini 3B
mistralai/voxtral-mini-3b-2507A 3B-parameter speech-language model built on the Ministral-3B backbone with an audio encoder for state-of-the-art audio understanding. Supports speech transcription, translation, audio Q&A, and voice-to-function calling across 8 languages.
Z-Image Turbo
tongyi/z-image-turboGlobal: Generate realistic images