Claude Opus 4.7

Claude Opus 4.7 is a non-reasoning model from Anthropic. 41 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.63.3 ±3.0
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.43/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$25 batch $12.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1483 (#12)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

61,128 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
62.5
CodingCode writing and repair.
67.7
ReasoningLogic problems and puzzles.
57.5
MultimodalTasks with images and text.
63.5
KnowledgeFacts and expert knowledge.
63.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
49.4
MathMath problems.
66.4

Results

41 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIMEMath96.3%58.616 Apr 2026Vals AI
LiveBench MathematicsMath92.9%72.0xhigh effort25 Jun 2026LiveBench
GPQA DiamondKnowledge90.2%61.51 Sept 2026Vals AI
MMLU ProKnowledge89.9%62.21 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge88.5%59.2Artificial Analysis
LiveBench ReasoningReasoning87.2%77.1xhigh effort25 Jun 2026LiveBench
OTIS Mock AIME 2024-2025Math86.7%60.0max effortEpoch AI
GPQA diamondKnowledge86.4%58.0max effortEpoch AI
MMMU ProMultimodal85.5%66.61 Sept 2026Vals AI
LiveCodeBenchCoding85.1%61.71 Sept 2026Vals AI
SWE-Bench verifiedCoding83.5%64.5max effortEpoch AI
React Native EvalsCoding82.8%63.3Callstack
LiveBench CodingCoding82.1%73.9xhigh effort25 Jun 2026LiveBench
SWE-benchCoding82.0%63.31 Sept 2026Vals AI
LiveBench Data AnalysisReasoning78.3%64.7xhigh effort25 Jun 2026LiveBench
LiveBench LanguageKnowledge77.9%64.6xhigh effort25 Jun 2026LiveBench
Artificial Analysis MMMU-ProMultimodal76.4%60.5Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning75.7%60.6Artificial Analysis
τ²-Bench Tool-Agent-User EvaluationAgentic74.0%51.9Victor Barres et al.
Vibe Code Bench v1.1Coding71.0%71.7OpenHands21 Sept 2026Vals AI
FrontierMath-Tiers-1-3-v2-PrivateMath70.2%70.9max effortEpoch AI
Terminal-Bench 2.0Agentic68.5%71.64 Jun 2026Vals AI
Terminal-Bench 2.1Agentic68.5%64.521 Sept 2026Vals AI
LiveBench Instruction FollowingInstruction66.7%65.4xhigh effort25 Jun 2026LiveBench
Gert Labs Composite Game BenchmarkAgentic65.6%70.6Gert Labs
SimpleQA VerifiedKnowledge51.7%69.4xhigh effortEpoch AI
LiveBench Agentic CodingAgentic50.7%64.4xhigh effort25 Jun 2026LiveBench
IOI v1Coding47.1%66.79 Aug 2026Vals AI
Artificial Analysis Omniscience AccuracyKnowledge44.7%69.2Artificial Analysis
Code MigrationCoding43.9%73.021 Sept 2026Vals AI
FrontierMath-2025-02-28-PrivateMath43.8%71.5xhigh effortEpoch AI
Artificial Analysis IFBenchInstruction43.6%33.4Artificial Analysis
τ²-bench BankingAgentic40.2%27.7max effort · Sierra4 Aug 2026Sierra Research
FrontierCode 1.1 MainCoding38.5%67.7Cognition
Furniture AssemblyReasoning33.3%63.2max effortEpoch AI
Artificial Analysis Humanity's Last ExamKnowledge33.3%60.7Artificial Analysis
FrontierMath-Tier-4-v2-PrivateMath31.7%63.2max effortEpoch AI
MirrorCodeCoding31.1%high effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge30.9%61.1Artificial Analysis
Mystery Game PuzzlesReasoning28.0%63.7max effortEpoch AI
FrontierMath-Tier-4-2025-07-01-PrivateMath22.9%68.3xhigh effortEpoch AI
ResearchClawBenchAgentic20.7%InternScience
EBR-benchReasoning19.0%62.4max effortEpoch AI
OSWorld 2.0Agentic13.9%62.6Mengqi Yuan et al.
Chess PuzzlesReasoning7.0%32.6max effortEpoch AI
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic7.0%64.0NeoCognition
Critical Physics TasksReasoning5.1%49.1Artificial Analysis
ProgramBenchCoding0.0%21 Sept 2026Vals AI

41 benchmarks count, from 45 of 48 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSierra Research, collected directlyMIT — results are in the licensed repository

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1