DeepSeek API: models and pricing

FastMetal serves 4 DeepSeek models through one OpenAI-compatible API key, billed per token from a prepaid balance. Change the base URL and key, and existing code calls them as is. The list also shows DeepSeek models we do not serve yet.

Sort by:

16 models

DeepSeek V4.1 Flash

deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...

Provider:deepseek logodeepseek
Pricing:$0.211 in · $0.633 out / 1M tokens
Context:1.0M tokens

DeepSeek V4 Flash

deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...

Provider:deepseek logodeepseek
Pricing:$0.095 in · $0.1899 out / 1M tokens
Context:1.0M tokens
Text Overall
#103

DeepSeek V4 Pro

deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...

Provider:deepseek logodeepseek
Pricing:$1.38 in · $2.75 out / 1M tokens
Context:1.0M tokens
Text Overall
#62

DeepSeek V3.2

deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#122

DeepSeek V3.2 Speciale

deepseek/deepseek-v3.2-speciale

DeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens

DeepSeek V3.2 Exp

deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#127

DeepSeek V3.1 Terminus

deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#138

DeepSeek V3.1 Terminus (exacto)

deepseek/deepseek-v3.1-terminus:exacto

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens

DeepSeek V3.1

deepseek/deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:33K tokens
Text Overall
#134

R1 0528

deepseek/deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#128

DeepSeek V3 0324

deepseek/deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#164

R1 Distill Qwen 32B

deepseek/deepseek-r1-distill-qwen-32b

DeepSeek R1 Distill Qwen 32B is a distilled large language model based on [Qwen 2.5 32B](https://huggingface.co/Qwen/Qwen2.5-32B), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:33K tokens

R1 Distill Llama 70B

deepseek/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:131K tokens

R1

deepseek/deepseek-r1

DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:64K tokens
Text Overall
#162

DeepSeek V3

deepseek/deepseek-chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions.

Provider:deepseek logodeepseek
Pricing:Not on FastMetal yet
Context:164K tokens
Text Overall
#205

deepseek-v4-flash-0731

deepseek-v4-flash-0731

Global: DeepSeek's V4 Flash 0731

Provider:deepseek logodeepseek
Pricing:$0.1477 in · $0.2954 out / 1M tokens
Context:- tokens