OpenAI
availableShows if the model has enough results for an index.GPT-5.6 Luna
GPT-5.6 Luna is a reasoning model from OpenAI in the GPT-5.6 family. 53 benchmarks count toward its score, in 7 categories.
IndexOverall score. 50 is the middle.69.9 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.58/s
Input / 1MUS dollars per 1M input tokens.$0.2 batch $0.1
Output / 1MUS dollars per 1M output tokens.$1.2 batch $0.6 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1430 (#86)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
28,547 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
53 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 98.3% | 66.5 | max effort | — | Epoch AI |
| SWE-bench | Coding | 93.0% | 72.1 | — | 1 Sept 2026 | Vals AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 92.3% | 63.5 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 92.3% | 63.5 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 91.7% | 62.9 | — | 1 Sept 2026 | Vals AI |
| GPQA diamond | Knowledge | 91.6% | 62.8 | max effort | — | Epoch AI |
| Artificial Analysis GPQA Diamond | Knowledge | 91.1% | 61.9 | — | — | Artificial Analysis |
| ARC-AGI-1 (semi-private) | Reasoning | 88.0% | 68.8 | max effort | — | ARC Prize Foundation |
| Artificial Analysis Harvey LAB-AA | Agentic | 87.9% | 69.8 | — | — | Artificial Analysis |
| MMLU Pro | Knowledge | 86.0% | 56.1 | — | 1 Sept 2026 | Vals AI |
| VulcanBench v3 | Coding | 85.5% | 70.3 | — | — | VulcanBench contributors |
| MMMU Pro | Multimodal | 85.0% | 65.7 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 83.7% | 66.1 | — | — | Artificial Analysis |
| BrowseComp | Agentic | 83.3% | 69.7 | — | — | OpenAI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 82.1% | 77.7 | max effort | — | Epoch AI |
| MMMU-Pro with Python | Multimodal | 79.5% | — | — | — | OpenAI |
| Terminal-Bench 2.1 | Agentic | 79.0% | 70.7 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis MMMU-Pro | Multimodal | 78.6% | 63.1 | — | — | Artificial Analysis |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 78.4% | 54.9 | — | — | MMMU-Pro authors |
| CyberGym | Agentic | 77.9% | 69.5 | — | — | Zhun Wang et al. |
| Vibe Code Bench v1.1 | Coding | 77.1% | 74.2 | OpenHands | 21 Sept 2026 | Vals AI |
| IOI v1 | Coding | 72.9% | 81.4 | — | 9 Aug 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 71.5% | 69.3 | — | — | Artificial Analysis |
| EuroEval French | Multilingual | 68.9% | 88.3 | — | — | EuroEval |
| DeepSWE | Agentic | 67.2% | 72.6 | — | — | Datacurve AI |
| EuroEval Swedish | Multilingual | 66.9% | 85.8 | — | — | EuroEval |
| SWE-bench Pro | Coding | 62.7% | 64.6 | — | — | Xiang Deng et al. |
| IOI | Coding | 61.8% | 72.6 | — | 21 Sept 2026 | Vals AI |
| cursorBench32 | Coding | 61.1% | 69.6 | — | — | Benchmark authors |
| FrontierMath-Tier-4-v2-Private | Math | 61.0% | 77.2 | max effort | — | Epoch AI |
| EuroEval Portuguese | Multilingual | 60.7% | 78.1 | — | — | EuroEval |
| SkillsBench | Coding | 60.4% | 74.2 | OpenHands | 11 Sept 2026 | Vals AI |
| EuroEval Italian | Multilingual | 60.0% | 77.2 | — | — | EuroEval |
| ProofBench v1.1 | Math | 60.0% | 73.7 | — | 21 Sept 2026 | Vals AI |
| ARC-AGI-2 (semi-private) | Reasoning | 59.5% | 69.7 | max effort | — | ARC Prize Foundation |
| EuroEval Dutch | Multilingual | 59.3% | 76.3 | — | — | EuroEval |
| EuroEval Spanish | Multilingual | 57.9% | 74.5 | — | — | EuroEval |
| HealthBench Professional | Knowledge | 55.7% | — | — | — | Rebecca Soskin Hicks et al. |
| FrontierCode 1.1 Extended | Coding | 55.1% | — | — | — | Cognition |
| EuroEval Polish | Multilingual | 54.2% | 69.9 | — | — | EuroEval |
| Artificial Analysis SciCode | Coding | 53.6% | 67.2 | — | — | Artificial Analysis |
| Toolathlon | Agentic | 53.4% | 67.8 | — | — | OpenAI |
| EuroEval German | Multilingual | 52.6% | 68.0 | — | — | EuroEval |
| Artificial Analysis Intelligence Index | Knowledge | 51.2% | 86.5 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 47.2% | 71.8 | — | — | Artificial Analysis |
| OSWorld 2.0 | Agentic | 45.6% | 77.7 | — | — | Mengqi Yuan et al. |
| Code Migration | Coding | 44.5% | 73.4 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 42.7% | 66.7 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 42.7% | 72.0 | — | — | Artificial Analysis |
| Furniture Assembly | Reasoning | 42.5% | 69.8 | max effort | — | Epoch AI |
| SimpleQA Verified | Knowledge | 41.0% | 59.5 | max effort | — | Epoch AI |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 40.8% | 63.2 | — | — | Artificial Analysis |
| Artificial Analysis ITBench-AA | Agentic | 40.3% | — | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 40.0% | 75.2 | max effort | — | Epoch AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 39.5% | 67.4 | — | — | Artificial Analysis |
| APEX-Agents-AA | Agentic | 35.8% | 69.3 | — | — | Artificial Analysis / Mercor |
| HealthBench Hard | Knowledge | 32.0% | 72.3 | — | — | Meta AI |
| Artificial Analysis Tau3-Banking | Agentic | 31.1% | 55.0 | — | — | Artificial Analysis |
| Artificial Analysis GDP.pdf | Agentic | 24.0% | 75.4 | — | — | Artificial Analysis |
| Vibe Code Bench 1-100 | Coding | 22.6% | 76.6 | OpenHands | 16 Sept 2026 | Vals AI |
| Mystery Game Puzzles | Reasoning | 21.0% | 56.3 | max effort | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 20.6% | 81.6 | — | — | Artificial Analysis |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 19.4% | 54.6 | — | — | Wisedocs and Artificial Analysis |
| Terminal-Bench 4.0.0 | Agentic | 17.3% | 71.6 | max effort · Codex | 21 Sept 2026 | Terminal-Bench |
| Terminal-Bench 3.0 | Agentic | 14.3% | 65.8 | — | — | Ryan Marten et al. |
| ExploitGym | Agentic | 12.4% | 70.1 | — | — | Zhun Wang et al. |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 7.0% | 64.0 | — | — | NeoCognition |
| Terminal-Bench 4.0 | Agentic | 4.5% | 62.7 | — | 21 Sept 2026 | Vals AI |
| Agent Arena command recovery | Agentic | 3.7 | 71.5 | xhigh effort | 15 Sept 2026 | LMArena |
| Agent Arena steerability | Agentic | 2.1 | 69.8 | xhigh effort | 15 Sept 2026 | LMArena |
| ARC-AGI-3 (semi-private) | Reasoning | 0.2% | — | max effort | — | ARC Prize Foundation |
| ProgramBench | Coding | 0.0% | — | — | 21 Sept 2026 | Vals AI |
| Agent Arena task outcome | Agentic | -5.0 | 61.7 | xhigh effort | 15 Sept 2026 | LMArena |
53 benchmarks count, from 67 of 73 results. A grey row does not count. Too few models took that benchmark.