Grok 4.7

Grok 4.7 is a reasoning model from xAI. 24 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.70.1 ±4.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.51/s
Input / 1MUS dollars per 1M input tokens.$1.6
Output / 1MUS dollars per 1M output tokens.$4.8
ContextMaximum tokens in one request.500K
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
66.6
CodingCode writing and repair.
71.9
ReasoningLogic problems and puzzles.
67.9
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
72.9
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
78.8
MathMath problems.
67.9

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
LiveBench MathematicsMath95.7%75.8xhigh effort25 Jun 2026LiveBench
Vibe Code Bench v1.1Coding86.2%78.0OpenHands21 Sept 2026Vals AI
LiveBench ReasoningReasoning82.7%70.8xhigh effort25 Jun 2026LiveBench
LiveBench LanguageKnowledge80.1%67.2xhigh effort25 Jun 2026LiveBench
LiveBench CodingCoding77.2%65.8xhigh effort25 Jun 2026LiveBench
LiveBench Data AnalysisReasoning76.9%62.8xhigh effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning76.7%61.3Artificial Analysis
LiveBench Instruction FollowingInstruction75.3%78.8xhigh effort25 Jun 2026LiveBench
Terminal-Bench 2.1Agentic73.4%67.421 Sept 2026Vals AI
DeepSWEAgentic71.0%75.4Datacurve AI
Artificial Analysis AutomationBenchAgentic65.6%79.3Artificial Analysis
EEBench V1 core corpusCoding64.0%atopile
GDPval-AA normalizedAgentic59.8%81.5Artificial Analysis
IOICoding57.7%70.821 Sept 2026Vals AI
Artificial Analysis SciCodeCoding57.4%72.5Artificial Analysis
HealthBench ProfessionalKnowledge56.7%Rebecca Soskin Hicks et al.
LiveBench Agentic CodingAgentic54.0%67.5xhigh effort25 Jun 2026LiveBench
Artificial Analysis Omniscience AccuracyKnowledge47.4%72.5Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge46.5%80.5Artificial Analysis
cursorBench40Coding46.3%Benchmark authors
Code MigrationCoding44.8%73.621 Sept 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge43.1%71.3Artificial Analysis
Terminal-Bench 4.0.0Agentic37.6%85.9xhigh effort · Grok Build21 Sept 2026Terminal-Bench
FrontierSWE v2Coding29.5%70.8Proximal
ProofBench v1.1Math26.0%59.921 Sept 2026Vals AI
Artificial Analysis GDP.pdfAgentic20.0%71.1Artificial Analysis
Artificial Analysis Harvey LAB-AAAgentic19.6%5.0Artificial Analysis
Critical Physics TasksReasoning17.7%75.5Artificial Analysis

24 benchmarks count, from 25 of 28 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AITerminal-Bench, collected directlyNo licence stated for the leaderboard. The harness repo is Apache-2.0

Same level, lower price

GPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87

More from xAI

Grok 4.671.8Grok 4.566.0Grok 4.356.0Grok 4.20 (Reasoning)56.9Grok 451.6