Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It supports long-context inference up to 256k tokens, extended from the previous 128k. This update improves agentic coding with higher accuracy and better generalization across scaffolds, and enhances frontend coding with more aesthetic and functional outputs for web, 3D, and related tasks. Kimi K2 is optimized for agentic capabilities, including advanced tool use, reasoning, and code synthesis. The model is trained with a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.
Evaluations
20
across 12 benchmarks
Latency
1.5s
Context length
262k
tokens
Cost
$0.60 · $2.50
input · output per 1M tokens
Key takeaways
Kimi K2 0905 is a large-scale Mixture-of-Experts (MoE) language model primarily designed for agentic tasks, excelling in agentic and frontend coding, reasoning, and tool use, with strong capabilities in long-context inference up to 256k tokens.
The model boasts 1 trillion total parameters (32 billion active per forward pass) and was trained using a novel stack incorporating the MuonClip optimizer for stable large-scale MoE training.
It demonstrates strong mathematical problem-solving skills, achieving 96.25% accuracy on MATH-500, surpassing human performance, and often providing detailed step-by-step explanations. However, it struggles with vague parameters, indirect reasoning, nuanced interpretations, complex multi-turn tasks, external tool usage, and occasional calculation errors.
A significant limitation identified is its complete failure on the Berkeley-Function-Calling-v3 dataset (0% accuracy) due to consistent 'Invalid model ID' errors, indicating an API interaction or model setup issue rather than a performance failure on the task itself.
The model's strengths include proficiency in basic to intermediate algebra, understanding various mathematical concepts (number theory, geometry, precalculus), and handling complex problem structures effectively, but its weaknesses are primarily occasional calculation errors, misinterpretations of problem constraints, and overly verbose derivations in some cases.