The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve relevance & readability. It's also better at working with uploaded files, providing deeper insights & more thorough responses. GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level of GPT-4 Turbo while being twice as fast and 50% more cost-effective. GPT-4o also offers improved performance in processing non-English languages and enhanced visual capabilities.
Evaluations
10
across 10 benchmarks
Latency
646ms
Context length
128k
tokens
Cost
$2.50 · $10
input · output per 1M tokens
Key takeaways
GPT-4o (2024-11-20) is designed for multimodal interaction, specifically supporting text and image inputs to generate text outputs. It aims to improve creative writing, file processing, and visual comprehension, delivering natural, engaging, and personalized responses.
This model is a proprietary Transformer-based architecture developed by OpenAI, featuring an extensive context length of 128,000 tokens.
The model exhibits strong capabilities in creative writing, processing uploaded files, and performing well in non-English languages. However, it struggles with complex mathematical reasoning, precise formatting requirements, and advanced calculation accuracy, particularly in highly constrained or intricate problem-solving scenarios.
In the MATH-500 dataset, GPT-4o achieved 73.2% accuracy, demonstrating excellent logical reasoning and strong performance in algebra, number theory, and prealgebra problems, particularly excelling in problems requiring step-by-step reasoning. It struggled with abstract counting and probability problems, complex multi-step problems, and maintaining strict mathematical formatting.
In the AIME 2025 dataset (advanced mathematics), GPT-4o achieved 0% accuracy, consistently failing to produce correct final answers despite providing detailed step-by-step solutions. Its primary weaknesses were computational errors, misinterpretation of constraints, and incorrect application of advanced mathematical principles.