GPT-5.4 mini

GPT-5.4 mini is a reasoning model from OpenAI in the GPT-5.4 family. 47 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.56.8 ±2.4
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.62/s
Input / 1MUS dollars per 1M input tokens.$0.75 batch $0.375
Output / 1MUS dollars per 1M output tokens.$4.5 batch $2.25 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.400K
EloLMArena rating and rank.1412 (#127)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

59,387 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
57.5
CodingCode writing and repair.
56.8
ReasoningLogic problems and puzzles.
54.0
MultimodalTasks with images and text.
55.4
KnowledgeFacts and expert knowledge.
55.7
MultilingualTasks in many languages.
80.9
InstructionTasks with strict rules in the prompt.
59.0
MathMath problems.
55.4

Results

47 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIMEMath95.6%58.316 Apr 2026Vals AI
τ²-Bench Tool-Agent-User EvaluationAgentic93.4%65.8Victor Barres et al.
OTIS Mock AIME 2024-2025Math88.9%61.2xhigh effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge88.0%59.5David Rein et al.
Artificial Analysis GPQA DiamondKnowledge87.5%58.2Artificial Analysis
GPQA diamondKnowledge86.9%58.4xhigh effortEpoch AI
MMLU ProKnowledge84.6%53.81 Sept 2026Vals AI
GPQA DiamondKnowledge83.1%54.91 Sept 2026Vals AI
LiveCodeBenchCoding81.5%58.41 Sept 2026Vals AI
MMMU ProMultimodal79.2%56.31 Sept 2026Vals AI
LiveBench MathematicsMath78.5%52.7xhigh effort25 Jun 2026LiveBench
MMMU-Pro with PythonMultimodal78.0%OpenAI
Artificial Analysis Long Context ReasoningReasoning77.0%61.5Artificial Analysis
Massive Multi-discipline Multimodal Understanding ProMultimodal76.6%52.0MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal73.3%56.7Artificial Analysis
Artificial Analysis IFBenchInstruction73.3%63.5Artificial Analysis
SWE-benchCoding73.0%56.11 Sept 2026Vals AI
OSWorld-VerifiedAgentic72.1%60.8Tianbao Xie et al.
LiveBench CodingCoding71.6%56.6xhigh effort25 Jun 2026LiveBench
LiveBench ReasoningReasoning71.3%55.0xhigh effort25 Jun 2026LiveBench
LiveBench LanguageKnowledge71.0%56.2xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning70.8%54.3xhigh effort25 Jun 2026LiveBench
EuroEval SwedishMultilingual68.9%88.3high effortEuroEval
EuroEval FrenchMultilingual67.4%86.4high effortEuroEval
EuroEval FrenchMultilingual67.3%86.3medium effortEuroEval
EuroEval SwedishMultilingual67.3%86.3medium effortEuroEval
EuroEval FrenchMultilingual67.1%86.1low effortEuroEval
EuroEval ItalianMultilingual66.2%84.9high effortEuroEval
EuroEval ItalianMultilingual65.8%84.5medium effortEuroEval
EuroEval SwedishMultilingual65.0%83.5low effortEuroEval
EuroEval PortugueseMultilingual64.7%83.1high effortEuroEval
EuroEval PortugueseMultilingual63.7%81.8medium effortEuroEval
ARC-AGI-1 (semi-private)Reasoning63.7%57.3xhigh effortARC Prize Foundation
EuroEval DutchMultilingual63.1%81.1high effortEuroEval
EuroEval DutchMultilingual63.0%81.0medium effortEuroEval
EuroEval SpanishMultilingual62.8%80.7high effortEuroEval
EuroEval PolishMultilingual62.5%80.3high effortEuroEval
EuroEval ItalianMultilingual62.5%80.3low effortEuroEval
EuroEval DutchMultilingual62.1%79.8low effortEuroEval
EuroEval SpanishMultilingual60.8%78.2medium effortEuroEval
EuroEval PortugueseMultilingual60.6%77.9low effortEuroEval
EuroEval PolishMultilingual60.2%77.5medium effortEuroEval
LiveBench Instruction FollowingInstruction59.8%54.6xhigh effort25 Jun 2026LiveBench
EuroEval PolishMultilingual58.8%75.7low effortEuroEval
EuroEval SpanishMultilingual57.9%74.6low effortEuroEval
EuroEval GermanMultilingual57.8%74.4high effortEuroEval
MCP AtlasAgentic57.7%51.7OpenAI
EuroEval GermanMultilingual56.3%72.6medium effortEuroEval
Artificial Analysis Coding IndexCoding56.1%58.5Artificial Analysis
Terminal-Bench 2.1Agentic54.7%56.321 Sept 2026Vals AI
EuroEval GermanMultilingual54.6%70.5low effortEuroEval
Artificial Analysis SciCodeCoding52.1%65.2Artificial Analysis
FrontierMath-Tiers-1-3-v2-PrivateMath51.2%60.3xhigh effortEpoch AI
Vibe Code Bench v1.1Coding48.0%62.1OpenHands21 Sept 2026Vals AI
Terminal-Bench 2.0Agentic44.9%54.84 Jun 2026Vals AI
ToolathlonAgentic42.9%58.0OpenAI
LiveBench Agentic CodingAgentic41.7%55.9xhigh effort25 Jun 2026LiveBench
Humanity's Last ExamKnowledge41.5%63.9Center for AI Safety et al.
Artificial Analysis Omniscience AccuracyKnowledge37.5%60.3Artificial Analysis
SimpleQA VerifiedKnowledge29.4%48.8high effortEpoch AI
FrontierMath-2025-02-28-PrivateMath28.3%57.0high effortEpoch AI
APEX-Agents-AAAgentic28.2%63.3Artificial Analysis / Mercor
Humanity's Last Exam without toolsKnowledge28.2%52.7OpenAI
Artificial Analysis Humanity's Last ExamKnowledge28.1%55.0Artificial Analysis
FrontierCode 1.1 MainCoding27.0%57.4Cognition
GDPval-AA normalizedAgentic25.0%54.8Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge24.1%52.5Artificial Analysis
Chess PuzzlesReasoning24.0%54.6xhigh effortEpoch AI
Artificial Analysis Agentic IndexAgentic19.6%53.2Artificial Analysis
ARC-AGI-2 (semi-private)Reasoning18.9%49.2xhigh effortARC Prize Foundation
Code MigrationCoding12.9%53.221 Sept 2026Vals AI
Critical Physics TasksReasoning10.0%59.3Artificial Analysis
FrontierMath-Tier-4-v2-PrivateMath9.8%52.6xhigh effortEpoch AI
Mystery Game PuzzlesReasoning7.0%41.4medium effortEpoch AI
IOI v1Coding6.4%43.69 Aug 2026Vals AI
FrontierMath-Tier-4-2025-07-01-PrivateMath2.1%45.7high effortEpoch AI
ProgramBenchCoding0.0%21 Sept 2026Vals AI

47 benchmarks count, from 75 of 77 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3