Back to Models
deepseek logo
deepseek/deepseek-r1-distill-llama-70b
Not Available

R1 Distill Llama 70B

DeepSeek R1 Distill Llama 70B is a distilled large language model based on [Llama-3.3-70B-Instruct](/meta-llama/llama-3.3-70b-instruct), using outputs from [DeepSeek R1](/deepseek/deepseek-r1). The model combines advanced distillation techniques to achieve high performance across multiple benchmarks, including: - AIME 2024 pass@1: 70.0 - MATH-500 pass@1: 94.5 - CodeForces Rating: 1633 The model leverages fine-tuning from DeepSeek R1's outputs, enabling competitive performance comparable to larger frontier models.

1/23/2025
131,072 tokens

Specifications

Modalities

Input
text
Output
text

Supported Parameters

frequency_penalty
include_reasoning
logit_bias
max_tokens
min_p
presence_penalty
reasoning
repetition_penalty
response_format
seed
stop
structured_outputs
temperature
top_k
top_p

Max Output Tokens

16,384

Frequently asked questions

Is R1 Distill Llama 70B available on FastMetal?
Not at the moment. DeepSeek V4.1 Flash, from the same lab, is available on the FastMetal API today.
What is the context window of R1 Distill Llama 70B?
131,072 tokens, shared between the prompt and the response.