Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique ability to switch seamlessly between a thinking mode for complex reasoning and a non-thinking mode for efficient dialogue ensures versatile, high-quality performance.
Significantly outperforming prior models like QwQ and Qwen2.5, Qwen3 delivers superior mathematics, coding, commonsense reasoning, creative writing, and interactive dialogue capabilities. The Qwen3-30B-A3B variant includes 30.5 billion parameters (3.3 billion activated), 48 layers, 128 experts (8 activated per task), and supports up to 131K token contexts with YaRN, setting a new standard among open-source models.
Evaluations
19
across 19 benchmarks
Latency
3m 21s
Context length
41k
tokens
Cost
$0.06 · $0.22
input · output per 1M tokens
Key takeaways
Qwen3 30B A3B is a Mixture-of-Experts (MoE) Large Language Model designed for reasoning, multilingual support, and advanced agent tasks, featuring a unique ability to switch between 'thinking mode' for complex reasoning and 'non-thinking mode' for efficient dialogue.
It is a Transformer-based MoE architecture with 30.5 billion parameters (3.3 billion activated), 48 layers, 128 experts (8 activated per task), and supports up to 131K token contexts using YaRN.
The model demonstrates strong capabilities in mathematics, coding, commonsense reasoning, creative writing, and interactive dialogue, achieving 95.16% accuracy on "MATH-500 (Algebra)".
Its primary limitations include struggles with nuanced reasoning, complex prompts, external knowledge retrieval, accurate numerical calculations beyond direct code execution, and long contexts, particularly in specialized domains like Biology/Medicine.
The model's worst performance was 5.63% accuracy on "Humanity's Last Exam (Biology/Medicine)", highlighting significant weaknesses in complex domain-specific knowledge and multi-step reasoning tasks.