MiniMax M2.1
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency. Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance. To avoid degrading this model's performance, MiniMax highly recommends preserving reasoning between turns. Learn more about using reasoning_details to pass back reasoning in our [docs](https://openrouter.ai/docs/use-cases/reasoning-tokens#preserving-reasoning-blocks).
Specifications
Modalities
Supported Parameters
Reasoning Configuration
- Default
- Thinking on
- Turning thinking off
- Not possible (always thinks)
Frequently asked questions
- Is MiniMax M2.1 available on FastMetal?
- Not at the moment. MiniMax M2.7, from the same lab, is available on the FastMetal API today.
- What is the context window of MiniMax M2.1?
- 196,608 tokens, shared between the prompt and the response.