Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable agentic execution across extended workflows. It is especially effective for asynchronous agent pipelines where tasks unfold over time - large codebases, multi-stage debugging, and end-to-end project orchestration.
Beyond coding, Opus 4.7 brings improved knowledge work capabilities - from drafting documents and building presentations to analyzing data. It maintains coherence across very long outputs and extended sessions, making it a strong default for tasks that require persistence, judgment, and follow-through.
Evaluations
13
across 12 benchmarks
Latency
1.4s
Context length
1M
tokens
Cost
$5 · $25
input · output per 1M tokens
Key takeaways
Claude Opus 4.7 is a proprietary Transformer-based model designed for long-running, asynchronous agent tasks, excelling in multi-step processes like multi-stage debugging and end-to-end project orchestration, and significantly enhancing knowledge work capabilities.
The model demonstrates strong logical reasoning, high accuracy (97.4%) for advanced mathematical problems (MATH-500), and maintains coherence across very long outputs and extended sessions, with a context length of 1,000,000 tokens and 128,000 max output tokens.
Known weaknesses include struggles with highly intricate enumeration, recursive counting, and precise value computation, leading to frequent calculation errors in complex multi-step mathematical problems and inconsistent performance in multimodal interpretation.
Despite its strength in complex math (MATH-500: 97.4% accuracy), it performs less robustly on broad, frontier-of-knowledge academic benchmarks like "Humanity's Last Exam" (30.8% accuracy), often due to numerical precision errors and misinterpretation of subtle constraints.
Key failure modes include misinterpretation of problem constraints, incorrect specialized knowledge in niche domains, flawed multimodal analysis, and logical gaps in complex deductions, alongside frequent arithmetic mistakes.