Back to Models
inception/mercury-2
Not Available

Mercury 2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving >1,000 tokens/sec on standard GPUs. Mercury 2 is 5x+ faster than leading speed-optimized LLMs like Claude 4.5 Haiku and GPT 5 Mini, at a fraction of the cost. Mercury 2 supports tunable reasoning levels, 128K context, native tool use, and schema-aligned JSON output. Built for coding workflows where latency compounds, real-time voice/search, and agent loops. OpenAI API compatible. Read more in the [blog post](https://www.inceptionlabs.ai/blog/introducing-mercury-2).

3/4/2026
128,000 tokens
#124 Code (Overall)

Specifications

Modalities

Input
text
Output
text

Supported Parameters

include_reasoning
max_tokens
reasoning
response_format
stop
structured_outputs
temperature
tool_choice
tools

Max Output Tokens

50,000

Reasoning Configuration

Default
Thinking on (effort: medium)
Selectable effort levels
high
medium
low
none
Turning thinking off
Possible (reasoning.enabled: false)

How to control thinking, and what it costs

Frequently asked questions

Is Mercury 2 available on FastMetal?
Not at the moment. Mercury 2.5, from the same lab, is available on the FastMetal API today.
What is the context window of Mercury 2?
128,000 tokens, shared between the prompt and the response.
How does Mercury 2 rank?
#209 on the public arena's Overall board (ELO 1,347). Ranks move as the leaderboard is updated.

Leaderboard

Text
OverallELO: 1,347
#209
ChineseELO: 1,402
#169
EnglishELO: 1,370
#204
russianELO: 1,304
#243
CodingELO: 1,394
#202
Creative WritingELO: 1,300
#224
Instruction FollowingELO: 1,324
#217
Hard PromptsELO: 1,360
#211
Multi-TurnELO: 1,343
#206
Code
OverallELO: 1,166
#124