Grok 4.6

Grok 4.6 is a reasoning model from xAI. 50 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.71.8 ±3.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.52/s
Input / 1MUS dollars per 1M input tokens.$2
Output / 1MUS dollars per 1M output tokens.$6
ContextMaximum tokens in one request.500K
EloLMArena rating and rank.1430 (#87)

50 is the middle of the board. The range shows the doubt in the index.

15,521 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
74.7
CodingCode writing and repair.
70.8
ReasoningLogic problems and puzzles.
69.2
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
69.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
73.4
MathMath problems.
68.1

Results

50 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math99.2%67.0xhigh effortEpoch AI
SWE-benchCoding95.6%74.21 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge94.9%65.8Artificial Analysis
GPQA DiamondKnowledge94.7%65.71 Sept 2026Vals AI
GPQA diamondKnowledge93.2%64.3xhigh effortEpoch AI
LiveBench MathematicsMath92.6%71.625 Jun 2026LiveBench
LiveBench ReasoningReasoning90.5%81.725 Jun 2026LiveBench
MMLU ProKnowledge89.4%61.41 Sept 2026Vals AI
LiveCodeBenchCoding88.2%64.51 Sept 2026Vals AI
VulcanBench v3Coding87.0%72.7VulcanBench contributors
ARC-AGI-1 (semi-private)Reasoning87.0%68.3xhigh effortARC Prize Foundation
LiveBench LanguageKnowledge83.7%71.525 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning80.3%63.8Artificial Analysis
Terminal-Bench 2.1Agentic78.3%70.321 Sept 2026Vals AI
Artificial Analysis Coding IndexCoding76.8%73.1Artificial Analysis
LiveBench CodingCoding76.8%65.125 Jun 2026LiveBench
Vibe Code Bench v1.1Coding76.2%73.9OpenHands21 Sept 2026Vals AI
LiveBench Data AnalysisReasoning73.9%58.525 Jun 2026LiveBench
LiveBench Instruction FollowingInstruction71.9%73.425 Jun 2026LiveBench
cursorBench32Coding70.8%78.6Benchmark authors
ARC-AGI-2 (semi-private)Reasoning67.1%73.6xhigh effortARC Prize Foundation
Artificial Analysis AutomationBenchAgentic66.7%80.8Artificial Analysis
FrontierMath-Tiers-1-3-v2-PrivateMath66.0%68.6xhigh effortEpoch AI
DeepSWEAgentic65.9%71.7Datacurve AI
FrontierCode 1.1 ExtendedCoding61.3%Cognition
APEX-AgentsAgentic57.5%87.1Moonshot AI / APEX-Agents benchmark authors
LiveBench Agentic CodingAgentic57.0%70.425 Jun 2026LiveBench
Artificial Analysis SciCodeCoding56.5%71.3Artificial Analysis
SkillsBenchCoding55.8%70.3OpenHands11 Sept 2026Vals AI
GDPval-AA normalizedAgentic55.3%78.0Artificial Analysis
Artificial Analysis Agentic IndexAgentic53.4%80.7Artificial Analysis
ProofBench v1.1Math51.0%70.121 Sept 2026Vals AI
Artificial Analysis Tau3-BankingAgentic50.7%81.4Artificial Analysis
SimpleQA VerifiedKnowledge48.9%66.8xhigh effortEpoch AI
Artificial Analysis EnterpriseOps-GymAgentic48.3%73.4Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge48.2%73.5Artificial Analysis
IOICoding47.6%66.421 Sept 2026Vals AI
Code MigrationCoding44.6%73.521 Sept 2026Vals AI
Artificial Analysis Intelligence IndexKnowledge44.3%77.8Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge42.9%71.1Artificial Analysis
Artificial Analysis AnalystAgentAgentic41.3%70.5Artificial Analysis
Mystery Game PuzzlesReasoning34.0%70.1xhigh effortEpoch AI
FrontierMath-Tier-4-v2-PrivateMath31.7%63.2xhigh effortEpoch AI
Chess PuzzlesReasoning31.0%63.6xhigh effortEpoch AI
EBR-benchReasoning30.5%70.2xhigh effortEpoch AI
Bug Hunt BenchCoding27.0%Pawel Huryn
Terminal-Bench 3.0Agentic26.5%74.9Ryan Marten et al.
FrontierSWE v2Coding25.3%68.5Proximal
Terminal-Bench 4.0.0Agentic20.3%73.8high effort · Grok Build21 Sept 2026Terminal-Bench
Terminal-Bench 4.0Agentic17.2%71.621 Sept 2026Vals AI
Critical Physics TasksReasoning17.1%74.2Artificial Analysis
Artificial Analysis GDP.pdfAgentic17.0%67.9Artificial Analysis
Vibe Code Bench 1-100Coding14.8%68.0OpenHands16 Sept 2026Vals AI
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic13.0%67.3NeoCognition
Agent Arena command recoveryAgentic6.674.8xhigh effort15 Sept 2026LMArena
Agent Arena steerabilityAgentic4.772.7xhigh effort15 Sept 2026LMArena
ARC-AGI-3 (semi-private)Reasoning2.1%xhigh effortARC Prize Foundation
Agent Arena task outcomeAgentic-2.564.5xhigh effort15 Sept 2026LMArena

50 benchmarks count, from 55 of 58 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useTerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0LMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87DeepSeek V4 Pro 081368.8 · $0.87

More from xAI

Grok 4.770.1Grok 4.566.0Grok 4.356.0Grok 4.20 (Reasoning)56.9Grok 451.6