MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1 to extend into general office work, reaching fluency in generating and operating Word, Excel, and Powerpoint files, context switching between diverse software environments, and working across different agent and human teams. MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency. Compared to its predecessor, M2.1 delivers cleaner, more concise outputs and faster perceived response times. It shows leading multilingual coding performance across major systems and application languages, achieving 49.4% on Multi-SWE-Bench and 72.5% on SWE-Bench Multilingual, and serves as a versatile agent “brain” for IDEs, coding tools, and general-purpose assistance. MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while maintaining low latency and deployment efficiency.
Evaluations
19
across 12 benchmarks
Latency
1m 4s
Context length
205k
tokens
Cost
$0.30 · $1.20
input · output per 1M tokens
Key takeaways
MiniMax M2.5 is a large language model designed for real-world productivity and general office work, extending its predecessor's coding capabilities to handle Word, Excel, and PowerPoint files and facilitate context-switching across diverse software environments and teams.
The model features an MoE (Mixture of Experts) architecture with 230 billion parameters, optimized for efficiency, low latency, and scalability, boasting a substantial context length of 204,800 tokens.
MiniMax M2.5 demonstrates high proficiency in mathematics, achieving 96.25% accuracy on the 'MATH-500' benchmark by providing detailed, step-by-step reasoning and strong calculation accuracy.
A significant limitation of the model is its very low accuracy of 14.63% on the 'Humanity's Last Exam' benchmark, indicating poor performance in quantitative and highly specific factual recall tasks across diverse specialized domains.
While excelling in detailed mathematical reasoning, the model struggles with complex multi-step reasoning, specialized knowledge in diverse domains, and long document reasoning in clinical notes, often making minor calculation errors or misinterpreting output formats.