Back to Models
inception/mercury-2
Not Available
Mercury 2
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs. Mercury 2 is 5x+ faster than leading speed-optimized LLMs like Claude 4.5 Haiku and GPT 5 Mini, at a fraction of the cost. Mercury 2 supports tunable reasoning levels, 128K context, native tool use, and schema-aligned JSON output. Built for coding workflows where latency compounds, real-time voice/search, and agent loops. OpenAI API compatible. Read more in the [blog post](https://www.inceptionlabs.ai/blog/introducing-mercury-2).
3/4/2026
128,000 tokens
#124 Code (Overall)
Specifications
Modalities
Input
text
Output
text
Supported Parameters
include_reasoning
max_tokens
reasoning
response_format
stop
structured_outputs
temperature
tool_choice
tools
Max Output Tokens
50,000Reasoning Configuration
- Default
- Thinking on (effort: medium)
- Selectable effort levels
- highmediumlownone
- Turning thinking off
- Possible (reasoning.enabled: false)
Frequently asked questions
- Is Mercury 2 available on FastMetal?
- Not at the moment. Mercury 2.5, from the same lab, is available on the FastMetal API today.
- What is the context window of Mercury 2?
- 128,000 tokens, shared between the prompt and the response.
- How does Mercury 2 rank?
- #209 on the public arena's Overall board (ELO 1,347). Ranks move as the leaderboard is updated.
Leaderboard
Text
OverallELO: 1,347
#209ChineseELO: 1,402
#169EnglishELO: 1,370
#204russianELO: 1,304
#243CodingELO: 1,394
#202Creative WritingELO: 1,300
#224Instruction FollowingELO: 1,324
#217Hard PromptsELO: 1,360
#211Multi-TurnELO: 1,343
#206Code
OverallELO: 1,166
#124