GPT-5.4 nano

GPT-5.4 nano is a reasoning model from OpenAI in the GPT-5.4 family. 46 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.53.1 ±2.4
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.76/s
Input / 1MUS dollars per 1M input tokens.$0.2 batch $0.1
Output / 1MUS dollars per 1M output tokens.$1.25 batch $0.625 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.400K
EloLMArena rating and rank.1373 (#167)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

58,424 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
52.2
CodingCode writing and repair.
55.3
ReasoningLogic problems and puzzles.
53.4
MultimodalTasks with images and text.
44.0
KnowledgeFacts and expert knowledge.
47.5
MultilingualTasks in many languages.
68.7
InstructionTasks with strict rules in the prompt.
66.1
MathMath problems.
57.2

Results

46 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic92.5%65.2Victor Barres et al.
LiveBench MathematicsMath91.0%69.5xhigh effort25 Jun 2026LiveBench
AIMEMath88.8%55.116 Apr 2026Vals AI
OTIS Mock AIME 2024-2025Math87.8%60.6high effortEpoch AI
LiveCodeBenchCoding84.0%60.71 Sept 2026Vals AI
Graduate-Level Google-Proof Q&AKnowledge82.8%54.7David Rein et al.
Artificial Analysis GPQA DiamondKnowledge81.7%52.3Artificial Analysis
LiveBench ReasoningReasoning81.1%68.6xhigh effort25 Jun 2026LiveBench
GPQA diamondKnowledge78.5%50.7high effortEpoch AI
GPQA DiamondKnowledge77.5%49.81 Sept 2026Vals AI
MMLU ProKnowledge77.2%42.11 Sept 2026Vals AI
Artificial Analysis Long Context ReasoningReasoning76.7%61.3Artificial Analysis
Artificial Analysis IFBenchInstruction75.9%66.1Artificial Analysis
MMMU ProMultimodal73.6%47.11 Sept 2026Vals AI
LiveBench CodingCoding70.8%55.3xhigh effort25 Jun 2026LiveBench
SWE-benchCoding69.8%53.51 Sept 2026Vals AI
MMMU-Pro with PythonMultimodal69.5%OpenAI
LiveBench Data AnalysisReasoning67.6%49.9xhigh effort25 Jun 2026LiveBench
LiveBench Instruction FollowingInstruction67.2%66.1xhigh effort25 Jun 2026LiveBench
Massive Multi-discipline Multimodal Understanding ProMultimodal66.1%34.9MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal65.4%47.0Artificial Analysis
LiveBench LanguageKnowledge62.5%46.1xhigh effort25 Jun 2026LiveBench
EuroEval SwedishMultilingual59.1%76.1high effortEuroEval
EuroEval SwedishMultilingual57.8%74.4medium effortEuroEval
EuroEval PortugueseMultilingual56.7%73.1high effortEuroEval
EuroEval SwedishMultilingual56.4%72.7low effortEuroEval
MCP AtlasAgentic56.1%50.5OpenAI
Artificial Analysis Coding IndexCoding56.1%58.5Artificial Analysis
EuroEval PortugueseMultilingual55.8%72.0medium effortEuroEval
EuroEval ItalianMultilingual55.7%71.9medium effortEuroEval
EuroEval PolishMultilingual55.2%71.3high effortEuroEval
EuroEval ItalianMultilingual55.1%71.0high effortEuroEval
EuroEval FrenchMultilingual54.9%70.8medium effortEuroEval
EuroEval FrenchMultilingual54.8%70.7high effortEuroEval
EuroEval SpanishMultilingual53.4%68.9medium effortEuroEval
EuroEval SpanishMultilingual53.3%68.8high effortEuroEval
EuroEval DutchMultilingual53.1%68.6medium effortEuroEval
EuroEval FrenchMultilingual53.0%68.5low effortEuroEval
EuroEval DutchMultilingual53.0%68.5high effortEuroEval
EuroEval PortugueseMultilingual52.7%68.1low effortEuroEval
EuroEval ItalianMultilingual52.5%67.9low effortEuroEval
EuroEval PolishMultilingual52.4%67.8medium effortEuroEval
EuroEval DutchMultilingual51.6%66.7low effortEuroEval
ARC-AGI-1 (semi-private)Reasoning51.5%51.5xhigh effortARC Prize Foundation
EuroEval SpanishMultilingual50.8%65.8low effortEuroEval
EuroEval PolishMultilingual50.7%65.6low effortEuroEval
EuroEval GermanMultilingual49.2%63.7medium effortEuroEval
EuroEval GermanMultilingual48.7%63.1high effortEuroEval
Artificial Analysis SciCodeCoding47.2%58.4Artificial Analysis
LiveBench Agentic CodingAgentic46.8%60.7xhigh effort25 Jun 2026LiveBench
EuroEval GermanMultilingual46.5%60.3low effortEuroEval
FrontierMath-Tiers-1-3-v2-PrivateMath44.9%56.7high effortEpoch AI
Terminal-Bench 2.1Agentic41.6%48.621 Sept 2026Vals AI
Terminal-Bench 2.0Agentic39.9%51.14 Jun 2026Vals AI
OSWorld-VerifiedAgentic39.0%29.9Tianbao Xie et al.
Humanity's Last ExamKnowledge37.7%60.7Center for AI Safety et al.
ToolathlonAgentic35.5%51.1OpenAI
Chess PuzzlesReasoning30.0%62.3high effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge28.3%55.2Artificial Analysis
Vibe Code Bench v1.1Coding26.1%53.0OpenHands21 Sept 2026Vals AI
FrontierMath-2025-02-28-PrivateMath25.9%54.8high effortEpoch AI
Artificial Analysis Omniscience AccuracyKnowledge25.7%45.7Artificial Analysis
APEX-Agents-AAAgentic24.9%60.7Artificial Analysis / Mercor
Humanity's Last Exam without toolsKnowledge24.3%49.4OpenAI
GDPval-AA normalizedAgentic21.8%52.3Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge20.7%48.3Artificial Analysis
Artificial Analysis Agentic IndexAgentic17.7%51.6Artificial Analysis
IOI v1Coding15.3%48.69 Aug 2026Vals AI
Code MigrationCoding14.5%54.221 Sept 2026Vals AI
FrontierMath-Tier-4-v2-PrivateMath12.2%53.8high effortEpoch AI
SimpleQA VerifiedKnowledge11.7%32.4high effortEpoch AI
Critical Physics TasksReasoning9.3%57.9Artificial Analysis
FrontierMath-Tier-4-2025-07-01-PrivateMath6.3%50.2high effortEpoch AI
ARC-AGI-2 (semi-private)Reasoning5.7%42.6xhigh effortARC Prize Foundation
Mystery Game PuzzlesReasoning5.0%39.2high effortEpoch AI

46 benchmarks count, from 74 of 75 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.572.3