Qwen API: models and pricing

FastMetal serves 8 Qwen models through one OpenAI-compatible API key, billed per token from a prepaid balance. Change the base URL and key, and existing code calls them as is. The list also shows Qwen models we do not serve yet.

Sort by:

42 models

Qwen3.8 Max Prime

qwen/qwen3.8-max-prime

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It accepts text, image, and video...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3.8 Max (0902)

qwen/qwen3.8-max-0902

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that accepts text, image, and video input and returns text,...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3.8 Flash

qwen/qwen3.8-flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document and codebase analysis, desktop interaction, chart analysis, and long-video analysis.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen3.8 27B

qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Provider:qwen logoqwen
Pricing:$0.211 in · $2.64 out / 1M tokens
Context:1.0M tokens
Text Overall
#99

Qwen3.8 27B (free)

qwen/qwen3.8-27b:free

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal interaction, and long-running agent tasks, with flexible thinking that can be...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3.8 2.4T A95B

qwen3.8-2.4t-a95b
ZDR

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...

Provider:qwen logoqwen
Pricing:$2.20 in · $6.60 out / 1M tokens
Context:1.0M tokens

Qwen3.8 Max

qwen3.8-max

Alibaba's flagship Qwen3.8 model: a 2.4T-parameter MoE (95B active) multimodal reasoning model with a 1M-token context window.

Provider:qwen logoqwen
Pricing:$2.11 in · $6.33 out / 1M tokens
Context:1.0M tokens
Text Overall
#23

Qwen3.7 Plus

qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#68

Qwen3.7 Max

qwen3.7-max
Policy unknown

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

Provider:qwen logoqwen
Pricing:$2.64 in · $7.92 out / 1M tokens
Context:262K tokens

Qwen3.6 Max Preview

qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#59

Qwen3.6 Plus

qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#88

Qwen3.5-122B-A10B

qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#135

Qwen3.5-27B

qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#150

Qwen3.5-35B-A3B

qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#168

Qwen3.5-Flash

qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens
Text Overall
#163

Qwen3.5 397B A17B

qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#92

Qwen3 Max

qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3 VL 235B A22B Instruct

qwen3-vl-235b-a22b-instruct

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video.

Provider:qwen logoqwen
Pricing:$0.211 in · $0.9284 out / 1M tokens
Context:262K tokens
Text Overall
#143

Qwen3 VL 235B A22B Thinking

qwen/qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#165

Qwen3 Next 80B A3B Instruct

qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.

Provider:qwen logoqwen
Pricing:$0.095 in · $1.17 out / 1M tokens
Context:262K tokens
Text Overall
#161

Qwen3 Next 80B A3B Instruct (free)

qwen/qwen3-next-80b-a3b-instruct:free

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3 Next 80B A3B Thinking

qwen/qwen3-next-80b-a3b-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:128K tokens
Text Overall
#195

Qwen3 30B A3B Thinking 2507

qwen/qwen3-30b-a3b-thinking-2507

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens

Qwen3 Coder 30B A3B Instruct

Qwen3-Coder-30B-A3B-Instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:160K tokens

Qwen3 30B A3B Instruct 2507

qwen/qwen3-30b-a3b-instruct-2507

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#185

Qwen3 235B A22B Thinking 2507

qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#160

Qwen3 Coder 480B A35B

qwen/qwen3-coder

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over repositories.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens

Qwen3 235B A22B Instruct 2507

qwen/qwen3-235b-a22b-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens
Text Overall
#126

Qwen3 235B A22B

qwen/qwen3-235b-a22b

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens
Text Overall
#189

Qwen3 30B A3B

qwen/qwen3-30b-a3b

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:41K tokens
Text Overall
#247

Qwen3 32B

qwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:41K tokens
Text Overall
#221

Qwen2.5 VL 32B Instruct

qwen/qwen2.5-vl-32b-instruct

Qwen2.5-VL-32B is a multimodal vision-language model fine-tuned through reinforcement learning for enhanced mathematical reasoning, structured outputs, and visual problem-solving capabilities.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:128K tokens

QwQ 32B

qwen/qwq-32b

QwQ is the reasoning model of the Qwen series. Compared with conventional instruction-tuned models, QwQ, which is capable of thinking and reasoning, can achieve significantly enhanced performance in downstream tasks, especially hard problem…

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#236

Qwen2.5 VL 72B Instruct

qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens

Qwen-Max

qwen/qwen-max

Qwen-Max, based on Qwen2.5, provides the best inference performance among [Qwen models](/qwen), especially for complex multi-step tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens

Qwen-Plus

qwen/qwen-plus

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:1.0M tokens

Qwen VL Max

qwen/qwen-vl-max

Qwen VL Max is a visual understanding model with 7500 tokens context length. It excels in delivering optimal performance for a broader spectrum of complex tasks.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:131K tokens

Qwen2.5 Coder 32B Instruct

qwen/qwen-2.5-coder-32b-instruct

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following improvements upon CodeQwen1.5: - Significantly improvements in **code generation**, **code reaso…

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#308

Qwen2.5 72B Instruct

qwen/qwen-2.5-72b-instruct

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge and has greatly improved capabilities in coding and mathematics, thanks to our specialized…

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#283

qwen3.6-27b

qwen3.6-27b

Global: Qwen 3.6 27B

Provider:qwen logoqwen
Pricing:$0.3376 in · $3.38 out / 1M tokens
Context:- tokens

Qwen3.6 35B-A3B

japan-qwen3.6-35b

Japan: Qwen 3.6 35B-A3B - hosted and run in Japan

Provider:qwen logoqwen
Pricing:$0.204 in · $1.01 out / 1M tokens
Context:- tokens

Qwen3 Coder 480B A35B

Qwen3-Coder-480B-A35B-Instruct-FP8

Alibaba's most capable open-source agentic coding model. A Mixture-of-Experts architecture with 480B total parameters (35B active), trained on 7.5 trillion tokens.

Provider:qwen logoqwen
Pricing:Not on FastMetal yet
Context:262K tokens