Back to Models
nvidia logo
nvidia/llama-3-1-nemotron-ultra-253b-v1
Not Available

Llama 3.1 Nemotron Ultra 253B v1

Llama-3.1-Nemotron-Ultra-253B-v1 is a large language model (LLM) optimized for advanced reasoning, human-interactive chat, retrieval-augmented generation (RAG), and tool-calling tasks. Derived from Meta’s Llama-3.1-405B-Instruct, it has been significantly customized using Neural Architecture Search (NAS), resulting in enhanced efficiency, reduced memory usage, and improved inference latency. The model supports a context length of up to 128K tokens and can operate efficiently on an 8x NVIDIA H100 node. Note: you must include `detailed thinking on` in the system prompt to enable reasoning. Please see [Usage Recommendations](https://huggingface.co/nvidia/Llama-3_1-Nemotron-Ultra-253B-v1#quick-start-and-usage-recommendations) for more.

4/8/2025
131,072 tokens

Specifications

Modalities

Input
text
Output
text

Supported Parameters

frequency_penalty
include_reasoning
max_tokens
presence_penalty
reasoning
repetition_penalty
response_format
structured_outputs
temperature
top_k
top_p

Frequently asked questions

Is Llama 3.1 Nemotron Ultra 253B v1 available on FastMetal?
Not at the moment. GLM 5.3 Flash, from the same lab, is available on the FastMetal API today.
What is the context window of Llama 3.1 Nemotron Ultra 253B v1?
131,072 tokens, shared between the prompt and the response.