Back to Models
inception/mercury
Not Available

Mercury

Mercury is the first diffusion large language model (dLLM). Applying a breakthrough discrete diffusion approach, the model runs 5-10x faster than even speed optimized models like GPT-4.1 Nano and Claude 3.5 Haiku while matching their performance. Mercury's speed enables developers to provide responsive user experiences, including with voice agents, search interfaces, and chatbots. Read more in the [blog post] (https://www.inceptionlabs.ai/blog/introducing-mercury) here.

6/26/2025
128,000 tokens
#228 Text (Coding)

Specifications

Modalities

Input
text
Output
text

Supported Parameters

frequency_penalty
max_tokens
presence_penalty
response_format
stop
structured_outputs
temperature
tool_choice
tools
top_k
top_p

Max Output Tokens

16,384

Frequently asked questions

Is Mercury available on FastMetal?
Not at the moment. Mercury 2.5, from the same lab, is available on the FastMetal API today.
What is the context window of Mercury?
128,000 tokens, shared between the prompt and the response.
How does Mercury rank?
#261 on the public arena's Overall board (ELO 1,308). Ranks move as the leaderboard is updated.

Leaderboard

Text
OverallELO: 1,308
#261
EnglishELO: 1,321
#267
CodingELO: 1,371
#228
Creative WritingELO: 1,224
#295
Instruction FollowingELO: 1,271
#279
Hard PromptsELO: 1,320
#254
Multi-TurnELO: 1,301
#252