Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated from final answers.
Compared to earlier Qwen3-30B releases, this version improves performance across logical reasoning, mathematics, science, coding, and multilingual benchmarks. It also demonstrates stronger instruction following, tool use, and alignment with human preferences. With higher reasoning efficiency and extended output budgets, it is best suited for advanced research, competitive problem solving, and agentic applications requiring structured long-context reasoning.
Evaluations
9
across 9 benchmarks
Latency
1m 44s
Context length
262k
tokens
Cost
$0.05 · $0.34
input · output per 1M tokens
Key takeaways
Qwen3 30B A3B Thinking 2507 is a 30B parameter Mixture-of-Experts (MoE) reasoning model optimized for complex multi-step thinking, primarily in a 'thinking mode' that separates internal reasoning from final answers, suitable for advanced research and agentic applications requiring structured long-context reasoning.
The model uses an MoE architecture with a substantial context length of 262,144 tokens, and it is an open-weights model with an open license.
The model excels in structured, rule-based logical reasoning tasks such as boolean algebra (100% accuracy), spatial navigation (94.4% accuracy), and multi-step arithmetic (96.5% accuracy). It also demonstrates strong instruction following and tool use, achieving 87.48% accuracy on 'Big Bench Hard'.
The model performs exceptionally well on direct mathematical problem-solving tasks, achieving high accuracy (93.3%) on the AIME 2024 benchmark, particularly in algebra, geometry, and number theory problems. Its main weakness lies in combinatorial and counting problems, where subtle interpretations of conditions or exhaustive enumeration can lead to errors.
The model struggles significantly with nuanced human-like causal judgment (65.5% accuracy), common-sense reasoning, processing implicit social rules in counterfactual scenarios, and performs poorly in 'movie_recommendation' tasks (58.3%) by over-relying on superficial attributes. It scored a very low 3% on 'Humanity's Last Exam', indicating severe challenges in complex reasoning across diverse specialized domains, particularly with quantitative, visual, and nuanced questions, and struggles with multi-step reasoning problems and interpreting specialized terminology.