Back to Models
inception/mercury
Not Available
Mercury
Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots. Read more in the [blog post] (https://www.inceptionlabs.ai/blog/introducing-mercury) here.
6/26/2025
128,000 tokens
#228 Text (Coding)
Specifications
Modalities
Input
text
Output
text
Supported Parameters
frequency_penalty
max_tokens
presence_penalty
response_format
stop
structured_outputs
temperature
tool_choice
tools
top_k
top_p
Max Output Tokens
16,384Frequently asked questions
- Is Mercury available on FastMetal?
- Not at the moment. Mercury 2.5, from the same lab, is available on the FastMetal API today.
- What is the context window of Mercury?
- 128,000 tokens, shared between the prompt and the response.
- How does Mercury rank?
- #261 on the public arena's Overall board (ELO 1,308). Ranks move as the leaderboard is updated.
Leaderboard
Text
OverallELO: 1,308
#261EnglishELO: 1,321
#267CodingELO: 1,371
#228Creative WritingELO: 1,224
#295Instruction FollowingELO: 1,271
#279Hard PromptsELO: 1,320
#254Multi-TurnELO: 1,301
#252