GPT-5.4 Pro

GPT-5.4 Pro is a reasoning model from OpenAI in the GPT-5.4 family. 12 benchmarks count toward its score, in 4 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.76.5 ±8.7
CoverageShare of the index weight with results.60%
SpeedOutput tokens per second.1/s
Input / 1MUS dollars per 1M input tokens.$30 batch $15
Output / 1MUS dollars per 1M output tokens.$180 batch $90 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
74.7
CodingCode writing and repair.
N/A
ReasoningLogic problems and puzzles.
85.9
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
67.2
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
78.8

Results

12 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
GPQA diamondKnowledge94.6%65.6xhigh effortEpoch AI
ARC-AGI-1 (semi-private)Reasoning94.5%71.8xhigh effortARC Prize Foundation
International Physics Olympiad 2025 (Theory)Math93.5%Meta AI
BrowseCompAgentic89.3%74.7OpenAI
ARC-AGI-2 (semi-private)Reasoning83.3%81.8xhigh effortARC Prize Foundation
FrontierMath-Tiers-1-3-v2-PrivateMath82.5%77.9xhigh effortEpoch AI
Humanity's Last ExamKnowledge58.7%78.5Center for AI Safety et al.
Chess PuzzlesReasoning58.6%95.0xhigh effortEpoch AI
FrontierMath-Tier-4-v2-PrivateMath58.5%76.0xhigh effortEpoch AI
FrontierMath-2025-02-28-PrivateMath50.0%77.3xhigh effortEpoch AI
SimpleQA VerifiedKnowledge46.3%64.4xhigh effortEpoch AI
Humanity's Last Exam without toolsKnowledge42.7%65.0OpenAI
FrontierMath-Tier-4-2025-07-01-PrivateMath37.5%84.2Epoch AI
FrontierScienceKnowledge36.7%OpenAI
FrontierScience ResearchKnowledge36.7%Meta AI
Critical Physics TasksReasoning30.0%95.0Artificial Analysis

12 benchmarks count, from 13 of 16 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

MiMo-V2.6-Pro73.5 · $0.87Muse Spark 1.375.0 · $4.25GPT-5.6 Sol76.7 · $10

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3