xAI
availableShows if the model has enough results for an index.Grok 4.6
Grok 4.6 is a reasoning model from xAI. 50 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.71.8 ±3.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.52/s
Input / 1MUS dollars per 1M input tokens.$2
Output / 1MUS dollars per 1M output tokens.$6
ContextMaximum tokens in one request.500K
EloLMArena rating and rank.1430 (#87)
50 is the middle of the board. The range shows the doubt in the index.
15,521 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
50 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| OTIS Mock AIME 2024-2025 | Math | 99.2% | 67.0 | xhigh effort | — | Epoch AI |
| SWE-bench | Coding | 95.6% | 74.2 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis GPQA Diamond | Knowledge | 94.9% | 65.8 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 94.7% | 65.7 | — | 1 Sept 2026 | Vals AI |
| GPQA diamond | Knowledge | 93.2% | 64.3 | xhigh effort | — | Epoch AI |
| LiveBench Mathematics | Math | 92.6% | 71.6 | — | 25 Jun 2026 | LiveBench |
| LiveBench Reasoning | Reasoning | 90.5% | 81.7 | — | 25 Jun 2026 | LiveBench |
| MMLU Pro | Knowledge | 89.4% | 61.4 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 88.2% | 64.5 | — | 1 Sept 2026 | Vals AI |
| VulcanBench v3 | Coding | 87.0% | 72.7 | — | — | VulcanBench contributors |
| ARC-AGI-1 (semi-private) | Reasoning | 87.0% | 68.3 | xhigh effort | — | ARC Prize Foundation |
| LiveBench Language | Knowledge | 83.7% | 71.5 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis Long Context Reasoning | Reasoning | 80.3% | 63.8 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 78.3% | 70.3 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 76.8% | 73.1 | — | — | Artificial Analysis |
| LiveBench Coding | Coding | 76.8% | 65.1 | — | 25 Jun 2026 | LiveBench |
| Vibe Code Bench v1.1 | Coding | 76.2% | 73.9 | OpenHands | 21 Sept 2026 | Vals AI |
| LiveBench Data Analysis | Reasoning | 73.9% | 58.5 | — | 25 Jun 2026 | LiveBench |
| LiveBench Instruction Following | Instruction | 71.9% | 73.4 | — | 25 Jun 2026 | LiveBench |
| cursorBench32 | Coding | 70.8% | 78.6 | — | — | Benchmark authors |
| ARC-AGI-2 (semi-private) | Reasoning | 67.1% | 73.6 | xhigh effort | — | ARC Prize Foundation |
| Artificial Analysis AutomationBench | Agentic | 66.7% | 80.8 | — | — | Artificial Analysis |
| FrontierMath-Tiers-1-3-v2-Private | Math | 66.0% | 68.6 | xhigh effort | — | Epoch AI |
| DeepSWE | Agentic | 65.9% | 71.7 | — | — | Datacurve AI |
| FrontierCode 1.1 Extended | Coding | 61.3% | — | — | — | Cognition |
| APEX-Agents | Agentic | 57.5% | 87.1 | — | — | Moonshot AI / APEX-Agents benchmark authors |
| LiveBench Agentic Coding | Agentic | 57.0% | 70.4 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis SciCode | Coding | 56.5% | 71.3 | — | — | Artificial Analysis |
| SkillsBench | Coding | 55.8% | 70.3 | OpenHands | 11 Sept 2026 | Vals AI |
| GDPval-AA normalized | Agentic | 55.3% | 78.0 | — | — | Artificial Analysis |
| Artificial Analysis Agentic Index | Agentic | 53.4% | 80.7 | — | — | Artificial Analysis |
| ProofBench v1.1 | Math | 51.0% | 70.1 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Tau3-Banking | Agentic | 50.7% | 81.4 | — | — | Artificial Analysis |
| SimpleQA Verified | Knowledge | 48.9% | 66.8 | xhigh effort | — | Epoch AI |
| Artificial Analysis EnterpriseOps-Gym | Agentic | 48.3% | 73.4 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 48.2% | 73.5 | — | — | Artificial Analysis |
| IOI | Coding | 47.6% | 66.4 | — | 21 Sept 2026 | Vals AI |
| Code Migration | Coding | 44.6% | 73.5 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Intelligence Index | Knowledge | 44.3% | 77.8 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 42.9% | 71.1 | — | — | Artificial Analysis |
| Artificial Analysis AnalystAgent | Agentic | 41.3% | 70.5 | — | — | Artificial Analysis |
| Mystery Game Puzzles | Reasoning | 34.0% | 70.1 | xhigh effort | — | Epoch AI |
| FrontierMath-Tier-4-v2-Private | Math | 31.7% | 63.2 | xhigh effort | — | Epoch AI |
| Chess Puzzles | Reasoning | 31.0% | 63.6 | xhigh effort | — | Epoch AI |
| EBR-bench | Reasoning | 30.5% | 70.2 | xhigh effort | — | Epoch AI |
| Bug Hunt Bench | Coding | 27.0% | — | — | — | Pawel Huryn |
| Terminal-Bench 3.0 | Agentic | 26.5% | 74.9 | — | — | Ryan Marten et al. |
| FrontierSWE v2 | Coding | 25.3% | 68.5 | — | — | Proximal |
| Terminal-Bench 4.0.0 | Agentic | 20.3% | 73.8 | high effort · Grok Build | 21 Sept 2026 | Terminal-Bench |
| Terminal-Bench 4.0 | Agentic | 17.2% | 71.6 | — | 21 Sept 2026 | Vals AI |
| Critical Physics Tasks | Reasoning | 17.1% | 74.2 | — | — | Artificial Analysis |
| Artificial Analysis GDP.pdf | Agentic | 17.0% | 67.9 | — | — | Artificial Analysis |
| Vibe Code Bench 1-100 | Coding | 14.8% | 68.0 | OpenHands | 16 Sept 2026 | Vals AI |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 13.0% | 67.3 | — | — | NeoCognition |
| Agent Arena command recovery | Agentic | 6.6 | 74.8 | xhigh effort | 15 Sept 2026 | LMArena |
| Agent Arena steerability | Agentic | 4.7 | 72.7 | xhigh effort | 15 Sept 2026 | LMArena |
| ARC-AGI-3 (semi-private) | Reasoning | 2.1% | — | xhigh effort | — | ARC Prize Foundation |
| Agent Arena task outcome | Agentic | -2.5 | 64.5 | xhigh effort | 15 Sept 2026 | LMArena |
50 benchmarks count, from 55 of 58 results. A grey row does not count. Too few models took that benchmark.