Claude Opus 4.6

Claude Opus 4.6 is a non-reasoning model from Anthropic. 40 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.58.9 ±3.1
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.33/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$25 batch $12.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1498 (#5)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

75,878 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
60.8
CodingCode writing and repair.
60.6
ReasoningLogic problems and puzzles.
53.1
MultimodalTasks with images and text.
55.2
KnowledgeFacts and expert knowledge.
60.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
34.4
MathMath problems.
65.8

Results

40 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIME25 first-party comparison snapshotMath99.8%Arcee AI
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge95.0%76.2Xiaoxuan Du et al.
Graduate-Level Google-Proof Q&AKnowledge91.3%62.5David Rein et al.
OTIS Mock AIME 2024-2025Math91.1%62.5max effortEpoch AI
GPQA DiamondKnowledge89.2%60.6David Rein et al.
MMLU-Pro first-party comparison snapshotKnowledge89.1%60.9Arcee AI
GPQA diamondKnowledge88.4%59.8max effortEpoch AI
τ²-Bench Tool-Agent-User EvaluationAgentic84.8%59.7Victor Barres et al.
React Native EvalsCoding84.1%65.1Callstack
Artificial Analysis GPQA DiamondKnowledge84.0%54.6Artificial Analysis
BrowseCompAgentic83.7%70.1OpenAI
ScreenSpot ProMultimodal83.1%65.1Kaixin Li et al.
Massive Multitask Language Understanding ProfessionalKnowledge82.0%49.7Yubo Wang et al.
Software Engineering Benchmark VerifiedCoding80.8%62.4Carlos E. Jimenez et al.
Massive Multi-discipline Multimodal Understanding ProMultimodal77.3%53.1MMMU-Pro authors
SWE-Bench verifiedCoding77.2%59.4Epoch AI
SWE-bench Verified (mini-swe-agent-v2)Coding75.6%58.2Arcee AI
SWE-bench VerifiedCoding75.6%58.2mini-SWE-agent1 Sept 2026SWE-bench team
DeepSearchQAAgentic73.7%56.5Meta AI
OSWorld-VerifiedAgentic72.7%61.4Tianbao Xie et al.
Artificial Analysis MMMU-ProMultimodal72.5%55.7Artificial Analysis
SWE-bench MultilingualCoding72.0%mini-SWE-agent2 Sept 2026SWE-bench team
LiveCodeBench ProCoding70.7%LiveCodeBench Pro authors
Claw-EvalAgentic70.4%69.0Bowen Ye et al.
Artificial Analysis Long Context ReasoningReasoning67.0%54.6Artificial Analysis
CyberGymAgentic66.6%61.7Zhun Wang et al.
FrontierMath-Tiers-1-3-v2-PrivateMath66.0%68.6max effortEpoch AI
SWE-RebenchCoding65.3%Nebius
MedXpertQA MultimodalMultimodal64.8%Meta AI
Gert Labs Composite Game BenchmarkAgentic61.9%67.3Gert Labs
Vibe Code Bench v1.1Coding57.6%66.1OpenHands21 Sept 2026Vals AI
SWE-bench ProCoding53.4%55.6Xiang Deng et al.
Humanity's Last ExamKnowledge53.0%73.7Center for AI Safety et al.
MedXpertQA TextKnowledge52.1%Meta AI
ERQAMultimodal51.6%47.1Qwen
SimpleQA VerifiedKnowledge47.0%65.1max effortEpoch AI
Artificial Analysis Omniscience AccuracyKnowledge45.8%70.5Artificial Analysis
Artificial Analysis IFBenchInstruction44.6%34.4Artificial Analysis
FrontierMath-2025-02-28-PrivateMath40.7%68.6max effortEpoch AI
Humanity's Last Exam without toolsKnowledge40.0%62.7OpenAI
JobBenchAgentic36.7%59.3Yuetai Li et al.
Furniture AssemblyReasoning28.3%59.7max effortEpoch AI
τ²-bench BankingAgentic27.3%18.4max effort · Sierra4 Aug 2026Sierra Research
FrontierCode 1.1 MainCoding26.9%57.4Cognition
FrontierMath-Tier-4-v2-PrivateMath26.8%60.8max effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge26.4%55.4Artificial Analysis
Mystery Game PuzzlesReasoning25.0%60.5max effortEpoch AI
FrontierMath-Tier-4-2025-07-01-PrivateMath22.9%68.3max effortEpoch AI
ResearchClawBenchAgentic19.9%InternScience
Artificial Analysis Humanity's Last ExamKnowledge19.1%45.3Artificial Analysis
HealthBench HardKnowledge14.8%49.9Meta AI
Chess PuzzlesReasoning14.0%41.6max effortEpoch AI
EBR-benchReasoning12.7%58.0max effortEpoch AI
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic5.0%62.9NeoCognition
Critical Physics TasksReasoning2.8%44.2Artificial Analysis

40 benchmarks count, from 48 of 55 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be citedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISierra Research, collected directlyMIT — results are in the licensed repository

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1