MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency.
Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance.
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency.
Evaluations
23
across 11 benchmarks
Latency
1m 16s
Context length
205k
tokens
Cost
$0.30 · $1.20
input · output per 1M tokens
Key takeaways
MiniMax M2.1 is designed as a lightweight, state-of-the-art large language model (10B activated parameters, 230B total via MoE) optimized for coding, agentic workflows, and modern application development.
The model demonstrates strong numerical reasoning and logical deduction, achieving 94.85% accuracy on Grade School Math 8K by correctly handling multi-step problems and applying mathematical operations.
Key limitations include struggling with complex SQL query fixing (17.14% pass rate on BIRD-CRITIC), misinterpreting nuanced phrasing like '3 times more' in math problems, and inconsistent performance on multi-modal scientific questions.
The model exhibits foundational understanding across scientific disciplines but struggles with precision, complex multi-step computations, and highly specialized factual recall, scoring approximately 25.8% on Humanity's Last Exam.
Despite its ability to generate detailed explanations, these often contain subtle errors in reasoning or calculation, and the model frequently assigns high confidence to incorrect answers, indicating a need for improved uncertainty estimation.