16 models
DeepSeek V4.1 Flash
deepseek-v4.1-flashDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED) architecture. It activates 8B parameters on input and 16B on...
DeepSeek V4 Flash
deepseek-v4-flashDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters, supporting a 1M-token context window. It is designed for fast inference and...
DeepSeek V4 Pro
deepseek-v4-proDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a 1M-token context window. It is designed for advanced reasoning, coding,...
DeepSeek V3.2
deepseek/deepseek-v3.2DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance.
DeepSeek V3.2 Speciale
deepseek/deepseek-v3.2-specialeDeepSeek-V3.2-Speciale is a high-compute variant of DeepSeek-V3.2 optimized for maximum reasoning and agentic performance.
DeepSeek V3.2 Exp
deepseek/deepseek-v3.2-expDeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures.
DeepSeek V3.1 Terminus
deepseek/deepseek-v3.1-terminusDeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…
DeepSeek V3.1 Terminus (exacto)
deepseek/deepseek-v3.1-terminus:exactoDeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further…
DeepSeek V3.1
deepseek/deepseek-chat-v3.1DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates.
R1 0528
deepseek/deepseek-r1-0528May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens.
DeepSeek V3 0324
deepseek/deepseek-chat-v3-0324DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team.
R1 Distill Qwen 32B
deepseek/deepseek-r1-distill-qwen-32bDeepSeek R1 Distill Qwen 32B is a distilled large language model based on [Qwen 2.5 32B](https://huggingface.co/Qwen/Qwen2.5-32B), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).
R1 Distill Llama 70B
deepseek/deepseek-r1-distill-llama-70bDeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1).
R1
deepseek/deepseek-r1DeepSeek R1 is here: Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active in an inference pass.
DeepSeek V3
deepseek/deepseek-chatDeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions.
deepseek-v4-flash-0731
deepseek-v4-flash-0731Global: DeepSeek's V4 Flash 0731