Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with improvements across system design, code security, and specification adherence. The model is designed for extended autonomous operation, maintaining task continuity across sessions and providing fact-based progress tracking.
Sonnet 4.5 also introduces stronger agentic capabilities, including improved tool orchestration, speculative parallel execution, and more efficient context and memory management. With enhanced context tracking and awareness of token usage across tool calls, it is particularly well-suited for multi-context and long-running workflows. Use cases span software engineering, cybersecurity, financial analysis, research agents, and other domains requiring sustained reasoning and tool use.
Evaluations
15
across 13 benchmarks
Latency
1.8s
Context length
200k
tokens
Cost
$3 · $15
input · output per 1M tokens
Key takeaways
Claude Sonnet 4.5 is optimized for real-world agents and coding workflows, specifically for autonomous operation in multi-context, long-running tasks within software engineering, cybersecurity, and financial analysis. It focuses on solving complex coding and agentic problems.
It is a proprietary Mixture of Experts (MoE) model with a context length of 200,000 tokens, designed for extended autonomous operation. Key architectural capabilities include tool orchestration, speculative parallel execution, and efficient context/memory management for advanced agentic behavior.
The model excels at accurately following detailed function specifications, consistently using contextual reasoning, breaking down complex problems, and is highly proficient in common list manipulations, string processing, and basic mathematical operations. However, it struggles with complex geometric problems, advanced numerical methods, arithmetic errors, misinterpreting mathematical constraints, and specific non-standard logical nuances.
Its best performance was on the "Numerical Reasoning Benchmark" with 95.24% accuracy, demonstrating strong adherence to function specifications, contextual reasoning, and proficiency in arithmetic and multi-step word problems. It showed strong foundational programming skills and accurately implemented common algorithmic patterns.
The model's worst performance was on the "Humanity's Last Exam (Biology/Medicine & Physics subsets)" dataset, scoring 6.81% accuracy, indicating significant weaknesses in complex mathematics, multimodal content (especially nuanced visual interpretation), and higher-order relativistic physics problems. Common failure modes included incorrect geometric setups, arithmetic errors, misinterpretation of constraints, and a failure to validate solutions, making its reliability variable for advanced math tasks.