Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for tasks like math, coding, and logical inference, and a "non-thinking" mode for faster, general-purpose conversation. The model demonstrates strong performance in instruction-following, agent tool use, creative writing, and multilingual tasks across 100+ languages and dialects. It natively handles 32K token contexts and can extend to 131K tokens using YaRN-based scaling.
Evaluations
40
across 26 benchmarks
Latency
—
Context length
41k
tokens
Cost
$0.05 · $0.20
input · output per 1M tokens
Key takeaways
Qwen3 32B is a 32.8B parameter Transformer-based causal language model optimized for complex reasoning, efficient dialogue, and multilingual tasks (100+ languages). It supports context lengths up to 131K tokens via YaRN scaling.
The model features a unique 'thinking' mode for mathematical, coding, and logical inference tasks, and a 'non-thinking' mode for faster, general-purpose conversations.
Qwen3 32B excels in instruction-following, agent tool use, and creative writing, evidenced by 99% accuracy on the 'AI2 Reasoning Challenge - Easy' dataset, demonstrating strong factual recall and reasoning in science domains.
The model performed poorly on the 'Humanity's Last Exam' benchmark (Chinese, complex reasoning), achieving only 15.38% accuracy due to difficulties with multi-step analytical problems, specialized domain knowledge, output formatting, and inconsistent reasoning.
Another significant weakness was observed on the 'SimpleQA' dataset (Art and factual questions), where it scored only 5.48% accuracy, indicating issues with factual accuracy, precise numerical data extraction, and inconsistent output formatting.