GPT-5.5

GPT-5.5 is a reasoning model from OpenAI. 60 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.72.3 ±2.6
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.53/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$30 batch $15 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1466 (#31)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

66,317 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
68.9
CodingCode writing and repair.
68.9
ReasoningLogic problems and puzzles.
80.2
MultimodalTasks with images and text.
65.5
KnowledgeFacts and expert knowledge.
71.4
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
68.9
MathMath problems.
77.7

Results

60 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math100.0%67.4xhigh effortEpoch AI
τ²-Bench Tool-Agent-User EvaluationAgentic98.0%69.1Victor Barres et al.
LiveBench MathematicsMath95.9%76.0xhigh effort25 Jun 2026LiveBench
ARC-AGI-1 (semi-private)Reasoning95.0%72.1xhigh effortARC Prize Foundation
GPQA diamondKnowledge94.0%65.0xhigh effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge93.6%64.7David Rein et al.
GPQA DiamondKnowledge93.6%64.7David Rein et al.
Artificial Analysis GPQA DiamondKnowledge93.5%64.4Artificial Analysis
GPQA DiamondKnowledge93.2%64.31 Sept 2026Vals AI
LiveBench ReasoningReasoning89.7%80.5xhigh effort25 Jun 2026LiveBench
MMMU ProMultimodal88.3%71.01 Sept 2026Vals AI
MMLU ProKnowledge88.1%59.41 Sept 2026Vals AI
OpenAI MRCR v2 8-needle 128K-256KReasoning87.5%OpenAI
LiveBench LanguageKnowledge87.4%75.9xhigh effort25 Jun 2026LiveBench
LiveCodeBenchCoding85.3%61.91 Sept 2026Vals AI
FrontierMath-Tiers-1-3-v2-PrivateMath85.3%79.4xhigh effortEpoch AI
ARC-AGI-2 (semi-private)Reasoning85.0%82.6xhigh effortARC Prize Foundation
React Native EvalsCoding84.7%65.9Callstack
BrowseCompAgentic84.4%70.6OpenAI
Artificial Analysis Long Context ReasoningReasoning84.3%66.6Artificial Analysis
MMMU-Pro with PythonMultimodal83.2%OpenAI
OpenAI MRCR v2 8-needle 64K-128KReasoning83.1%OpenAI
SWE-benchCoding82.6%63.81 Sept 2026Vals AI
LiveBench CodingCoding82.1%74.0xhigh effort25 Jun 2026LiveBench
CyberGymAgentic81.8%72.1Zhun Wang et al.
LiveBench Data AnalysisReasoning81.6%69.3xhigh effort25 Jun 2026LiveBench
Massive Multi-discipline Multimodal Understanding ProMultimodal81.2%59.5MMMU-Pro authors
SWE-Bench verifiedCoding80.6%62.2xhigh effortEpoch AI
Artificial Analysis MMMU-ProMultimodal79.9%64.7Artificial Analysis
OSWorld-VerifiedAgentic78.7%67.0Tianbao Xie et al.
Terminal-Bench 2.1Agentic76.4%69.121 Sept 2026Vals AI
Artificial Analysis IFBenchInstruction75.9%66.1Artificial Analysis
MCP AtlasAgentic75.3%64.9OpenAI
Artificial Analysis Coding IndexCoding74.9%71.7Artificial Analysis
Terminal-Bench 2.0Agentic73.2%74.94 Jun 2026Vals AI
Gert Labs Composite Game BenchmarkAgentic72.9%77.1Gert Labs
FrontierMath-Tier-4-v2-PrivateMath72.5%82.7xhigh effortEpoch AI
LiveBench Instruction FollowingInstruction70.7%71.7xhigh effort25 Jun 2026LiveBench
Vibe Code Bench v1.1Coding69.8%71.2OpenHands21 Sept 2026Vals AI
SimpleQA VerifiedKnowledge63.0%79.9xhigh effortEpoch AI
SkillsBenchCoding62.2%75.7OpenHands11 Sept 2026Vals AI
cursorBench31Coding59.2%Benchmark authors
SWE-bench ProCoding58.6%60.7Xiang Deng et al.
cursorBench32Coding58.4%67.0Benchmark authors
Artificial Analysis Omniscience AccuracyKnowledge58.0%85.6Artificial Analysis
Mystery Game PuzzlesReasoning56.0%93.5xhigh effortEpoch AI
Artificial Analysis SciCodeCoding55.8%70.3Artificial Analysis
ToolathlonAgentic55.6%69.8OpenAI
OfficeQA ProMultimodal54.1%66.4OfficeQA Pro authors
Chess PuzzlesReasoning54.0%93.3xhigh effortEpoch AI
LiveBench Agentic CodingAgentic54.0%67.5xhigh effort25 Jun 2026LiveBench
Humanity's Last ExamKnowledge52.2%73.0Center for AI Safety et al.
FrontierMath-2025-02-28-PrivateMath51.7%78.8xhigh effortEpoch AI
Artificial Analysis AnalystAgentAgentic50.0%76.6Artificial Analysis
Artificial Analysis ITBench-AAAgentic45.8%Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge45.8%74.2Artificial Analysis
Code MigrationCoding45.2%73.821 Sept 2026Vals AI
τ²-bench BankingAgentic44.6%30.8xhigh effort · Sierra4 Aug 2026Sierra Research
Furniture AssemblyReasoning44.2%70.9xhigh effortEpoch AI
FrontierCode 1.1 MainCoding43.0%71.8Cognition
JobBenchAgentic42.7%63.5Yuetai Li et al.
GDPval-AA normalizedAgentic41.8%67.7Artificial Analysis
Humanity's Last Exam without toolsKnowledge41.4%63.9OpenAI
Artificial Analysis Intelligence IndexKnowledge38.4%70.4Artificial Analysis
APEX-Agents-AAAgentic37.7%70.8Artificial Analysis / Mercor
Artificial Analysis Agentic IndexAgentic37.3%67.6Artificial Analysis
OEIS Open LiteMath36.0%medium effortEpoch AI
FrontierMath-Tier-4-2025-07-01-PrivateMath35.4%81.9xhigh effortEpoch AI
EBR-benchReasoning34.3%72.8xhigh effortEpoch AI
Critical Physics TasksReasoning27.1%95.0Artificial Analysis
OEIS OpenMath26.2%medium effortEpoch AI
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic20.0%71.2NeoCognition
ResearchClawBenchAgentic17.0%InternScience
ExploitGymAgentic13.4%70.8Zhun Wang et al.
OSWorld 2.0Agentic13.0%62.2Mengqi Yuan et al.
MirrorCodeCoding10.0%high effortEpoch AI
Agent Arena command recoveryAgentic9.678.2xhigh effort15 Sept 2026LMArena
Agent Arena steerabilityAgentic6.674.8xhigh effort15 Sept 2026LMArena
ProgramBenchCoding0.5%21 Sept 2026Vals AI
ARC-AGI-3 (semi-private)Reasoning0.4%high effortARC Prize Foundation
FrontierMath-ErdosMath0.0%xhigh effortEpoch AI
Agent Arena task outcomeAgentic-0.666.7xhigh effort15 Sept 2026LMArena

60 benchmarks count, from 70 of 82 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISierra Research, collected directlyMIT — results are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from OpenAI

GPT-6 Astra83.5GPT-5.6 Sol76.7GPT-5.6 Terra73.4GPT-6 Sol76.9GPT-5.5 Pro77.9