Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It delivers higher accuracy in math, coding, logic, and science tasks, follows complex instructions in Chinese and English more reliably, reduces hallucinations, and produces higher-quality responses for open-ended Q&A, writing, and conversation. The model supports over 100 languages with stronger translation and commonsense reasoning, and is optimized for retrieval-augmented generation (RAG) and tool calling, though it does not include a dedicated “thinking” mode.
Evaluations
19
across 17 benchmarks
Latency
1.1s
Context length
256k
tokens
Cost
$1.20 · $6
input · output per 1M tokens
Key takeaways
Qwen3 Max is designed for advanced reasoning, instruction following, multilingual support, and open-ended Q&A, aiming to reduce hallucinations, improve response quality, and optimize for Retrieval-Augmented Generation (RAG) and tool calling.
Qwen3 Max is an MoE (Mixture of Experts) architecture model with a 256k context length, optimized for over 100 languages with stronger translation and commonsense reasoning, and excels at following complex instructions in both Chinese and English.
The model demonstrates significant weaknesses in coding challenges, particularly search algorithms and nuanced code interpretations, as evidenced by very poor performance on the Terminal-Bench Dataset (3.75% accuracy). It also lacks a dedicated 'thinking' mode.
The best performance was observed on the MATH-500 dataset with an accuracy of 96.2%, showcasing its strength in mathematical reasoning.
This model is a proprietary release from Qwen, an updated version of the Qwen3 series, with 175 parameters and a text-to-text modality.