Kimi K2.5 builds on Kimi K2 with continued pretraining over approximately 15T mixed visual and text tokens. Built as a native multimodal model, K2.5 delivers state-of-the-art coding and vision capabilities and a self-directed agent swarm paradigm.
For complex tasks, Kimi K2.5 can self-direct an agent swarm with up to 100 sub-agents, executing parallel workflows across up to 1,500 tool calls. Compared with a single-agent setup, this reduces execution time by up to 4.5x. The agent swarm is automatically created and orchestrated by Kimi K2.5 without any predefined subagents or workflow.
Kimi K2.5 builds on Kimi K2 with continued pretraining over approximately 15T mixed visual and text tokens. Built as a native multimodal model, K2.5 delivers state-of-the-art coding and vision capabilities and a self-directed agent swarm paradigm.
For complex tasks, Kimi K2.5 can self-direct an agent swarm with up to 100 sub-agents, executing parallel workflows across up to 1,500 tool calls. Compared with a single-agent setup, this reduces execution time by up to 4.5x. The agent swarm is automatically created and orchestrated by Kimi K2.5 without any predefined subagents or workflow.
Evaluations
30
across 14 benchmarks
Latency
4m 40s
Context length
262k
tokens
Cost
$0.60 · $3
input · output per 1M tokens
Key takeaways
Kimi K2.5 is a multimodal (text+image-to-text) Mixture of Experts (MoE) model, pretrained on 15T mixed visual and text tokens, designed for complex tasks with state-of-the-art coding and vision capabilities.
The model features a self-directed agent swarm that can orchestrate up to 100 sub-agents and 1,500 tool calls, reducing execution time by up to 4.5x compared to single-agent setups for complex tasks.
Kimi K2.5 demonstrates high accuracy (97.8%) and truthfulness (98.39%) on mathematical tasks (MATH-500, TruthfulQA), providing detailed step-by-step reasoning and adhering to formatting rules, although it may have minor calculation or formatting errors.
A significant limitation is its complete failure (0% success rate) on complex multi-round agentic tasks (SWE-agent benchmark) due to an inability to interpret problem descriptions, formulate actionable plans, and effectively use tools beyond basic file viewing.
Common failure modes include getting stuck in a 'viewing' phase, lack of targeted actions for problem-solving, repetitive responses, and difficulties with nuanced instructions, advanced SQL, and intricate logical requirements (31.33% accuracy on BIRD-CRITIC).