Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation. Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.
Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use.
Evaluations
10
across 10 benchmarks
Latency
1.1s
Context length
1M
tokens
Cost
$3 · $15
input · output per 1M tokens
Key takeaways
Claude Sonnet 4.6 is Anthropic's most advanced Sonnet-class model, optimized for real-world agents and coding workflows. It excels at iterative development, complex codebase navigation, and end-to-end project management.
The model utilizes a Mixture of Experts (MoE) architecture, handles text and image inputs (outputting text), and boasts a substantial context length of 1,000,000 tokens with a maximum of 128,000 output tokens.
Capabilities include strong foundational understanding and advanced problem-solving, particularly excelling in Algebra and Number Theory with nearly perfect accuracy (94.94%). It also demonstrates proficiency in scripting and system configuration.
Limitations include struggles with tasks requiring deep domain-specific knowledge, complex debugging in non-standard environments (e.g., Linux kernel compilation, advanced ML implementations), specialized external tool interaction, and low-level C programming.
Evaluation results show 94.94% accuracy on MATH-500, with best performance in Algebra and Number Theory. Worst performance includes misinterpreting problem constraints in Geometry, errors in algebraic simplification, and vulnerability to miscalculation in Counting & Probability tasks.