OpenAI
availableShows if the model has enough results for an index.GPT-5.5
GPT-5.5 is a reasoning model from OpenAI. 60 benchmarks count toward its score, in 7 categories.
IndexOverall score. 50 is the middle.72.3 ±2.6
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.53/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$30 batch $15 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1466 (#31)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
66,317 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
60 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 100.0% | 67.4 | xhigh effort | — | Epoch AI |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 98.0% | 69.1 | — | — | Victor Barres et al. |
| LiveBench Mathematics | Math | 95.9% | 76.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| ARC-AGI-1 (semi-private) | Reasoning | 95.0% | 72.1 | xhigh effort | — | ARC Prize Foundation |
| GPQA diamond | Knowledge | 94.0% | 65.0 | xhigh effort | — | Epoch AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 93.6% | 64.7 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 93.6% | 64.7 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 93.5% | 64.4 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 93.2% | 64.3 | — | 1 Sept 2026 | Vals AI |
| LiveBench Reasoning | Reasoning | 89.7% | 80.5 | xhigh effort | 25 Jun 2026 | LiveBench |
| MMMU Pro | Multimodal | 88.3% | 71.0 | — | 1 Sept 2026 | Vals AI |
| MMLU Pro | Knowledge | 88.1% | 59.4 | — | 1 Sept 2026 | Vals AI |
| OpenAI MRCR v2 8-needle 128K-256K | Reasoning | 87.5% | — | — | — | OpenAI |
| LiveBench Language | Knowledge | 87.4% | 75.9 | xhigh effort | 25 Jun 2026 | LiveBench |
| LiveCodeBench | Coding | 85.3% | 61.9 | — | 1 Sept 2026 | Vals AI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 85.3% | 79.4 | xhigh effort | — | Epoch AI |
| ARC-AGI-2 (semi-private) | Reasoning | 85.0% | 82.6 | xhigh effort | — | ARC Prize Foundation |
| React Native Evals | Coding | 84.7% | 65.9 | — | — | Callstack |
| BrowseComp | Agentic | 84.4% | 70.6 | — | — | OpenAI |
| Artificial Analysis Long Context Reasoning | Reasoning | 84.3% | 66.6 | — | — | Artificial Analysis |
| MMMU-Pro with Python | Multimodal | 83.2% | — | — | — | OpenAI |
| OpenAI MRCR v2 8-needle 64K-128K | Reasoning | 83.1% | — | — | — | OpenAI |
| SWE-bench | Coding | 82.6% | 63.8 | — | 1 Sept 2026 | Vals AI |
| LiveBench Coding | Coding | 82.1% | 74.0 | xhigh effort | 25 Jun 2026 | LiveBench |
| CyberGym | Agentic | 81.8% | 72.1 | — | — | Zhun Wang et al. |
| LiveBench Data Analysis | Reasoning | 81.6% | 69.3 | xhigh effort | 25 Jun 2026 | LiveBench |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 81.2% | 59.5 | — | — | MMMU-Pro authors |
| SWE-Bench verified | Coding | 80.6% | 62.2 | xhigh effort | — | Epoch AI |
| Artificial Analysis MMMU-Pro | Multimodal | 79.9% | 64.7 | — | — | Artificial Analysis |
| OSWorld-Verified | Agentic | 78.7% | 67.0 | — | — | Tianbao Xie et al. |
| Terminal-Bench 2.1 | Agentic | 76.4% | 69.1 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis IFBench | Instruction | 75.9% | 66.1 | — | — | Artificial Analysis |
| MCP Atlas | Agentic | 75.3% | 64.9 | — | — | OpenAI |
| Artificial Analysis Coding Index | Coding | 74.9% | 71.7 | — | — | Artificial Analysis |
| Terminal-Bench 2.0 | Agentic | 73.2% | 74.9 | — | 4 Jun 2026 | Vals AI |
| Gert Labs Composite Game Benchmark | Agentic | 72.9% | 77.1 | — | — | Gert Labs |
| FrontierMath-Tier-4-v2-Private | Math | 72.5% | 82.7 | xhigh effort | — | Epoch AI |
| LiveBench Instruction Following | Instruction | 70.7% | 71.7 | xhigh effort | 25 Jun 2026 | LiveBench |
| Vibe Code Bench v1.1 | Coding | 69.8% | 71.2 | OpenHands | 21 Sept 2026 | Vals AI |
| SimpleQA Verified | Knowledge | 63.0% | 79.9 | xhigh effort | — | Epoch AI |
| SkillsBench | Coding | 62.2% | 75.7 | OpenHands | 11 Sept 2026 | Vals AI |
| cursorBench31 | Coding | 59.2% | — | — | — | Benchmark authors |
| SWE-bench Pro | Coding | 58.6% | 60.7 | — | — | Xiang Deng et al. |
| cursorBench32 | Coding | 58.4% | 67.0 | — | — | Benchmark authors |
| Artificial Analysis Omniscience Accuracy | Knowledge | 58.0% | 85.6 | — | — | Artificial Analysis |
| Mystery Game Puzzles | Reasoning | 56.0% | 93.5 | xhigh effort | — | Epoch AI |
| Artificial Analysis SciCode | Coding | 55.8% | 70.3 | — | — | Artificial Analysis |
| Toolathlon | Agentic | 55.6% | 69.8 | — | — | OpenAI |
| OfficeQA Pro | Multimodal | 54.1% | 66.4 | — | — | OfficeQA Pro authors |
| Chess Puzzles | Reasoning | 54.0% | 93.3 | xhigh effort | — | Epoch AI |
| LiveBench Agentic Coding | Agentic | 54.0% | 67.5 | xhigh effort | 25 Jun 2026 | LiveBench |
| Humanity's Last Exam | Knowledge | 52.2% | 73.0 | — | — | Center for AI Safety et al. |
| FrontierMath-2025-02-28-Private | Math | 51.7% | 78.8 | xhigh effort | — | Epoch AI |
| Artificial Analysis AnalystAgent | Agentic | 50.0% | 76.6 | — | — | Artificial Analysis |
| Artificial Analysis ITBench-AA | Agentic | 45.8% | — | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 45.8% | 74.2 | — | — | Artificial Analysis |
| Code Migration | Coding | 45.2% | 73.8 | — | 21 Sept 2026 | Vals AI |
| τ²-bench Banking | Agentic | 44.6% | 30.8 | xhigh effort · Sierra | 4 Aug 2026 | Sierra Research |
| Furniture Assembly | Reasoning | 44.2% | 70.9 | xhigh effort | — | Epoch AI |
| FrontierCode 1.1 Main | Coding | 43.0% | 71.8 | — | — | Cognition |
| JobBench | Agentic | 42.7% | 63.5 | — | — | Yuetai Li et al. |
| GDPval-AA normalized | Agentic | 41.8% | 67.7 | — | — | Artificial Analysis |
| Humanity's Last Exam without tools | Knowledge | 41.4% | 63.9 | — | — | OpenAI |
| Artificial Analysis Intelligence Index | Knowledge | 38.4% | 70.4 | — | — | Artificial Analysis |
| APEX-Agents-AA | Agentic | 37.7% | 70.8 | — | — | Artificial Analysis / Mercor |
| Artificial Analysis Agentic Index | Agentic | 37.3% | 67.6 | — | — | Artificial Analysis |
| OEIS Open Lite | Math | 36.0% | — | medium effort | — | Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 35.4% | 81.9 | xhigh effort | — | Epoch AI |
| EBR-bench | Reasoning | 34.3% | 72.8 | xhigh effort | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 27.1% | 95.0 | — | — | Artificial Analysis |
| OEIS Open | Math | 26.2% | — | medium effort | — | Epoch AI |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 20.0% | 71.2 | — | — | NeoCognition |
| ResearchClawBench | Agentic | 17.0% | — | — | — | InternScience |
| ExploitGym | Agentic | 13.4% | 70.8 | — | — | Zhun Wang et al. |
| OSWorld 2.0 | Agentic | 13.0% | 62.2 | — | — | Mengqi Yuan et al. |
| MirrorCode | Coding | 10.0% | — | high effort | — | Epoch AI |
| Agent Arena command recovery | Agentic | 9.6 | 78.2 | xhigh effort | 15 Sept 2026 | LMArena |
| Agent Arena steerability | Agentic | 6.6 | 74.8 | xhigh effort | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 0.5% | — | — | 21 Sept 2026 | Vals AI |
| ARC-AGI-3 (semi-private) | Reasoning | 0.4% | — | high effort | — | ARC Prize Foundation |
| FrontierMath-Erdos | Math | 0.0% | — | xhigh effort | — | Epoch AI |
| Agent Arena task outcome | Agentic | -0.6 | 66.7 | xhigh effort | 15 Sept 2026 | LMArena |
60 benchmarks count, from 70 of 82 results. A grey row does not count. Too few models took that benchmark.