OpenAI
availableShows if the model has enough results for an index.GPT-5.6 Terra
GPT-5.6 Terra is a reasoning model from OpenAI in the GPT-5.6 family. 51 benchmarks count toward its score, in 8 categories.
IndexOverall score. 50 is the middle.73.4 ±2.3
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.36/s
Input / 1MUS dollars per 1M input tokens.$2 batch $1
Output / 1MUS dollars per 1M output tokens.$12 batch $6 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1446 (#52)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
28,119 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
51 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 99.7% | 67.3 | max effort | — | Epoch AI |
| ARC-AGI-1 (semi-private) | Reasoning | 96.5% | 72.8 | max effort | — | ARC Prize Foundation |
| SWE-bench | Coding | 95.4% | 74.1 | — | 1 Sept 2026 | Vals AI |
| GPQA diamond | Knowledge | 93.3% | 64.4 | max effort | — | Epoch AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 92.9% | 64.0 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 92.9% | 64.0 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 92.5% | 63.3 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 90.9% | 62.2 | — | 1 Sept 2026 | Vals AI |
| IOI | Coding | 87.6% | 83.8 | — | 21 Sept 2026 | Vals AI |
| BrowseComp | Agentic | 87.5% | 73.2 | — | — | OpenAI |
| VulcanBench v3 | Coding | 87.0% | 72.7 | — | — | VulcanBench contributors |
| MMLU Pro | Knowledge | 86.7% | 57.1 | — | 1 Sept 2026 | Vals AI |
| MMMU Pro | Multimodal | 86.5% | 68.1 | — | 1 Sept 2026 | Vals AI |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 86.3% | 60.7 | — | — | Victor Barres et al. |
| FrontierMath-Tiers-1-3-v2-Private | Math | 86.0% | 79.8 | max effort | — | Epoch AI |
| LiveCodeBench | Coding | 85.9% | 62.4 | — | 1 Sept 2026 | Vals AI |
| ARC-AGI-2 (semi-private) | Reasoning | 83.9% | 82.0 | max effort | — | ARC Prize Foundation |
| Artificial Analysis Long Context Reasoning | Reasoning | 83.0% | 65.7 | — | — | Artificial Analysis |
| MMMU-Pro with Python | Multimodal | 82.0% | — | — | — | OpenAI |
| CyberGym | Agentic | 81.8% | 72.1 | — | — | Zhun Wang et al. |
| LABBench2: An Improved Benchmark for AI Systems Performing Biology Research | Knowledge | 81.2% | — | — | — | Jon M. Laurent et al. |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 80.7% | 58.7 | — | — | MMMU-Pro authors |
| Artificial Analysis MMMU-Pro | Multimodal | 80.7% | 65.7 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 77.5% | 69.8 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 76.7% | 73.0 | — | — | Artificial Analysis |
| Vibe Code Bench v1.1 | Coding | 74.6% | 73.2 | OpenHands | 21 Sept 2026 | Vals AI |
| ProofBench v1.1 | Math | 74.0% | 79.4 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis IFBench | Instruction | 71.2% | 61.3 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-v2-Private | Math | 70.7% | 81.9 | max effort | — | Epoch AI |
| EuroEval Swedish | Multilingual | 70.6% | 90.4 | — | — | EuroEval |
| DeepSWE | Agentic | 69.6% | 74.4 | — | — | Datacurve AI |
| EuroEval French | Multilingual | 68.9% | 88.3 | — | — | EuroEval |
| EuroEval Italian | Multilingual | 67.8% | 87.0 | — | — | EuroEval |
| EuroEval Portuguese | Multilingual | 65.4% | 83.9 | — | — | EuroEval |
| IOI v1 | Coding | 65.3% | 77.0 | — | 9 Aug 2026 | Vals AI |
| EuroEval Dutch | Multilingual | 65.0% | 83.4 | — | — | EuroEval |
| cursorBench32 | Coding | 64.9% | 73.1 | — | — | Benchmark authors |
| SWE-bench Pro | Coding | 63.4% | 65.3 | — | — | Xiang Deng et al. |
| EuroEval Spanish | Multilingual | 62.1% | 79.8 | — | — | EuroEval |
| EuroEval German | Multilingual | 59.7% | 76.9 | — | — | EuroEval |
| EuroEval Polish | Multilingual | 59.0% | 75.9 | — | — | EuroEval |
| SkillsBench | Coding | 58.9% | 72.9 | OpenHands | 11 Sept 2026 | Vals AI |
| HealthBench Professional | Knowledge | 57.7% | — | — | — | Rebecca Soskin Hicks et al. |
| FrontierCode 1.1 Extended | Coding | 55.8% | — | — | — | Cognition |
| Artificial Analysis SciCode | Coding | 55.0% | 69.2 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 55.0% | 91.1 | — | — | Artificial Analysis |
| Furniture Assembly | Reasoning | 54.2% | 78.1 | max effort | — | Epoch AI |
| Chess Puzzles | Reasoning | 54.0% | 93.3 | max effort | — | Epoch AI |
| Toolathlon | Agentic | 53.1% | 67.5 | — | — | OpenAI |
| HLE-Verified | Knowledge | 51.1% | — | — | — | Weiqi Zhai et al. |
| Artificial Analysis ITBench-AA | Agentic | 51.0% | — | — | — | Artificial Analysis |
| OSWorld 2.0 | Agentic | 50.2% | 79.8 | — | — | Mengqi Yuan et al. |
| Code Migration | Coding | 47.8% | 75.5 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 46.8% | 71.8 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 46.6% | 71.4 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 43.7% | 72.8 | — | — | Artificial Analysis |
| SimpleQA Verified | Knowledge | 43.2% | 61.5 | max effort | — | Epoch AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 42.9% | 71.1 | — | — | Artificial Analysis |
| APEX-Agents-AA | Agentic | 38.9% | 71.8 | — | — | Artificial Analysis / Mercor |
| Mystery Game Puzzles | Reasoning | 35.0% | 71.1 | max effort | — | Epoch AI |
| HealthBench Hard | Knowledge | 32.7% | 73.2 | — | — | Meta AI |
| Critical Physics Tasks | Reasoning | 30.0% | 95.0 | — | — | Artificial Analysis |
| Terminal-Bench 4.0 | Agentic | 26.3% | 78.0 | — | 21 Sept 2026 | Vals AI |
| ExploitGym | Agentic | 23.2% | 77.8 | — | — | Zhun Wang et al. |
| Terminal-Bench 4.0.0 | Agentic | 21.5% | 74.6 | max effort · Codex | 21 Sept 2026 | Terminal-Bench |
| Terminal-Bench 3.0 | Agentic | 20.8% | 70.6 | — | — | Ryan Marten et al. |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 16.0% | 69.0 | — | — | NeoCognition |
| Vibe Code Bench 1-100 | Coding | 14.8% | 68.1 | OpenHands | 16 Sept 2026 | Vals AI |
| Agent Arena steerability | Agentic | 4.9 | 72.9 | xhigh effort | 15 Sept 2026 | LMArena |
| Agent Arena command recovery | Agentic | 3.0 | 70.8 | xhigh effort | 15 Sept 2026 | LMArena |
| ARC-AGI-3 (semi-private) | Reasoning | 0.8% | — | max effort | — | ARC Prize Foundation |
| ProgramBench | Coding | 0.5% | — | — | 21 Sept 2026 | Vals AI |
| Agent Arena task outcome | Agentic | -4.4 | 62.4 | xhigh effort | 15 Sept 2026 | LMArena |
51 benchmarks count, from 65 of 73 results. A grey row does not count. Too few models took that benchmark.