GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding, and tool use, while reducing latency and cost for large-scale deployments.
The model is designed for production environments that require a balance of capability and efficiency, making it well suited for chat applications, coding assistants, and agent workflows that operate at scale. GPT-5.4 mini delivers reliable instruction following, solid multi-step reasoning, and consistent performance across diverse tasks with improved cost efficiency.
Evaluations
18
across 18 benchmarks
Latency
444ms
Context length
400k
tokens
Cost
$0.75 · $4.50
input · output per 1M tokens
Key takeaways
GPT-5.4 Mini is designed for high-throughput, latency-sensitive applications such as chat, coding assistants, and agent workflows, prioritizing efficiency and cost-reduction for large-scale deployments through a Mixture of Experts (MoE) architecture.
It is a versatile model supporting both text and image inputs, excelling in reasoning, coding, and tool use, with strong performance in basic to intermediate mathematical problem-solving (86.6% accuracy on MATH-500 dataset, specifically algebra, number theory, and geometry).
A significant limitation is its overall low accuracy (3.33% on AIME 2024), struggling with multi-step calculations, numerical precision, open-ended scientific questions, and image interpretation.
Common failure modes include misinterpretation of problem statements, calculation errors in complex arithmetic, incorrect application of mathematical theorems, and logical errors, often coupled with overconfidence in incorrect answers.
The model demonstrates strengths in foundational mathematical understanding, particularly in Algebra and Number Theory, but its performance significantly declines with increased complexity and subtle logical conditions found in advanced mathematical problem-solving.