DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.
Evaluations
16
across 9 benchmarks
Latency
726ms
Context length
128k
tokens
Cost
$0.28 · $0.42
input · output per 1M tokens
Key takeaways
DeepSeek V3.2 is an 800-parameter Transformer-based LLM optimized for high computational efficiency, with strong reasoning and agentic tool-use capabilities.
It employs DeepSeek Sparse Attention (DSA) to reduce training and inference costs, demonstrating gold-medal results on the 2025 IMO and IOI.
The model excels in core mathematical domains like algebra, geometry, and number theory, and follows complex formatting instructions well.
Its limitations include struggles with complex combinatorial reasoning, subtle geometric interpretations, highly specific numerical calculations, and precise application of obscure definitions.
DeepSeek V3.2 achieved 91.4% accuracy on the MATH-500 dataset, demonstrating strong mathematical reasoning, but only 5.52% on the 'Humanity's Last Exam' truthfulness benchmark, indicating a significant gap in its ability to handle nuanced truthfulness across diverse domains.