Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built on the trillion-parameter Mixture-of-Experts (MoE) architecture introduced in Kimi K2, it activates 32 billion parameters per forward pass and supports 256 k-token context windows. The model is optimized for persistent step-by-step thought, dynamic tool invocation, and complex reasoning workflows that span hundreds of turns. It interleaves step-by-step reasoning with tool use, enabling autonomous research, coding, and writing that can persist for hundreds of sequential actions without drift.
Evaluations
20
across 9 benchmarks
Latency
—
Context length
262k
tokens
Cost
$0.60 · $2.50
input · output per 1M tokens
Key takeaways
Kimi K2 Thinking is designed for agentic, long-horizon reasoning, persistent step-by-step thinking, dynamic tool invocation, and complex reasoning workflows, enabling autonomous research, coding, and writing.
It is built on a trillion-parameter Mixture-of-Experts (MoE) architecture, activating 32 billion parameters per forward pass and supporting 256k-token context windows.
The model demonstrates strong factual recall and reasoning in MMLU tasks, particularly in algebra and number theory, achieving 94.6% accuracy with detailed step-by-step explanations.
A significant limitation is its 0% accuracy on the Terminal-Bench (Terminus-1) truthfulness benchmark, failing all 50 tasks due to difficulties in complex problem-solving, code generation, environment interaction, and adherence to specific output requirements.
Areas for improvement include precise geometric interpretation, rigorous adherence to output formatting, handling nuanced problem constraints, and a complete failure in multimodal understanding and image processing tasks.