DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism that reduces training and inference cost while preserving quality in long-context scenarios. A scalable reinforcement learning post-training framework further improves reasoning, with reported performance in the GPT-5 class, and the model has demonstrated gold-medal results on the 2025 IMO and IOI. V3.2 also uses a large-scale agentic task synthesis pipeline to better integrate reasoning into tool-use settings, boosting compliance and generalization in interactive environments.
Evaluations
16
across 9 benchmarks
Latency
726ms
Context length
128k
tokens
Cost
$0.28 · $0.42
input · output per 1M tokens
Key takeaways
DeepSeek V3.2 is an 800-parameter Transformer-based LLM optimized for high computational efficiency, strong reasoning, and agentic tool-use, designed to reduce training and inference costs with DeepSeek Sparse Attention (DSA).
The model excels in core mathematical domains like algebra, geometry, and number theory, demonstrating robust problem-solving (91.4% accuracy on MATH-500) and adherence to complex formatting instructions.
Despite strong reasoning, DeepSeek V3.2 struggles with complex combinatorial reasoning, subtle geometric interpretations, highly specific numerical calculations, and precise application of obscure definitions.
A significant limitation is its low truthfulness score (5.52% on 'Humanity's Last Exam'), indicating a gap in handling nuanced truthfulness across diverse domains and a tendency for numerical errors in multi-step calculations.
Its performance is sensitive to precise interpretation of problem statements and diagram cues, suggesting a need for improved contextual understanding in ambiguous scenarios.