Claude Opus 4.5

Claude Opus 4.5 is a non-reasoning model from Anthropic. 53 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.52.1 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.40/s
Input / 1MUS dollars per 1M input tokens.$5 batch $2.5
Output / 1MUS dollars per 1M output tokens.$25 batch $12.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.200K
EloLMArena rating and rank.1450 (#44)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

70,013 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
55.4
CodingCode writing and repair.
56.2
ReasoningLogic problems and puzzles.
51.0
MultimodalTasks with images and text.
39.2
KnowledgeFacts and expert knowledge.
55.4
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
37.8
MathMath problems.
51.7

Results

53 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MMLU-ReduxKnowledge96.6%56.5Qwen
AIME 2026Math95.1%54.3Qwen
MGSMMultilingual94.8%9 Jan 2026Vals AI
Harvard-MIT Mathematics Tournament November 2025Math93.3%Qwen
Harvard-MIT Mathematics Tournament February 2025Math92.9%51.6Qwen
τ²-bench TelecomAgentic92.3%65.1high effort · Sierra2 Mar 2026Sierra Research
C-EvalKnowledge92.2%C-Eval authors
Instruction-Following EvalInstruction90.9%47.1Jeffrey Zhou et al.
Massive Multitask Language Understanding ProfessionalKnowledge89.5%61.6Yubo Wang et al.
Artificial Analysis MMLU-ProKnowledge88.9%Artificial Analysis
Graduate-Level Google-Proof Q&AKnowledge87.0%58.6David Rein et al.
τ²-Bench Tool-Agent-User EvaluationAgentic86.3%60.7Victor Barres et al.
MMLU-ProXMultilingual85.7%MMLU-ProX authors
MMLU ProKnowledge85.6%55.41 Sept 2026Vals AI
GPQA diamondKnowledge85.5%57.2Epoch AI
Harvard-MIT Mathematics Tournament February 2026Math85.3%54.0Qwen
LiveCodeBench v6Coding84.8%54.2LiveCodeBench maintainers
VideoMMMUMultimodal84.4%Qwen
MMAnswerBenchMath84.0%Qwen
τ²-bench AirlineAgentic84.0%59.1high effort · Sierra2 Mar 2026Sierra Research
OTIS Mock AIME 2024-2025Math81.7%57.2Epoch AI
MMMU ProMultimodal81.1%59.31 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge81.0%51.6Artificial Analysis
Software Engineering Benchmark VerifiedCoding80.9%62.4Carlos E. Jimenez et al.
τ²-bench RetailAgentic79.6%55.9high effort · Sierra30 Apr 2026Sierra Research
GPQA DiamondKnowledge79.5%51.71 Sept 2026Vals AI
AIMEMath76.9%49.616 Apr 2026Vals AI
SWE-bench VerifiedCoding76.8%59.1medium effort · live-SWE-agent1 Sept 2026SWE-bench team
SWE-Bench verifiedCoding76.7%59.0Epoch AI
WideResearchAgentic76.4%62.0Qwen
LiveCodeBenchCoding75.0%52.51 Sept 2026Vals AI
MathVisionMultimodal74.3%Qwen
AI-NeedleReasoning74.0%Qwen
MCP-TasksAgentic71.8%Qwen
Artificial Analysis MMMU-ProMultimodal71.2%54.1Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning70.7%57.2Artificial Analysis
SWE-bench MultilingualCoding70.7%mini-SWE-agent2 Sept 2026SWE-bench team
Massive Multi-discipline Multimodal Understanding ProMultimodal70.6%42.2MMMU-Pro authors
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge70.6%55.9Xiaoxuan Du et al.
τ³-Bench Tool-Agent-User EvaluationAgentic70.2%54.8Sierra Research
CharXiv ReasoningMultimodal68.5%44.1CharXiv authors
V*Multimodal67.0%21.0Z.AI
OSWorld-VerifiedAgentic66.3%55.4Tianbao Xie et al.
OSWorldAgentic66.3%Z.AI
LongBench v2Reasoning64.4%LongBench v2 authors
Gert Labs Composite Game BenchmarkAgentic64.2%69.4Gert Labs
Claw-EvalAgentic59.6%52.1Bowen Ye et al.
Terminal-Bench 2.0Agentic58.4%64.44 Jun 2026Vals AI
Instruction Following BenchmarkInstruction58.0%33.6Benchmark authors
SWE-bench ProCoding57.1%59.2Xiang Deng et al.
NOVA-63Multilingual56.7%Qwen
Terminal-Bench 1.0Agentic56.3%59.212 Jan 2026Vals AI
SWE-bench (full test split)Coding52.6%Sonar Foundation Agent19 Dec 2025SWE-bench team
QwenClawBenchAgentic52.3%51.7Qwen
CyberGymAgentic50.6%50.8Zhun Wang et al.
ScreenSpot ProMultimodal45.7%26.2Kaixin Li et al.
SimpleQA VerifiedKnowledge45.7%63.9Epoch AI
ToolathlonAgentic43.5%58.5OpenAI
NL2RepoCoding43.2%58.8MiniMax
Artificial Analysis IFBenchInstruction43.0%32.7Artificial Analysis
MCP AtlasAgentic42.3%40.1OpenAI
Artificial Analysis Omniscience AccuracyKnowledge40.9%64.5Artificial Analysis
FrontierMath-Tiers-1-3-v2-PrivateMath34.4%50.8Epoch AI
JobBenchAgentic32.3%56.3Yuetai Li et al.
Humanity's Last ExamKnowledge30.8%54.9Center for AI Safety et al.
Furniture AssemblyReasoning28.3%59.7Epoch AI
DeepPlanningAgentic26.4%DeepPlanning authors
τ²-bench BankingAgentic24.7%16.6high effort · Sierra4 Aug 2026Sierra Research
Artificial Analysis Intelligence IndexKnowledge23.7%52.0Artificial Analysis
IOI v1Coding23.6%53.39 Aug 2026Vals AI
VITA-BenchAgentic23.3%42.2Meituan LongCat Team
Mystery Game PuzzlesReasoning22.0%57.3Epoch AI
FrontierMath-2025-02-28-PrivateMath20.7%49.9Epoch AI
EBR-benchReasoning14.3%59.1Epoch AI
Artificial Analysis Humanity's Last ExamKnowledge13.2%38.9Artificial Analysis
Chess PuzzlesReasoning8.0%33.9Epoch AI
FrontierMath-Tier-4-v2-PrivateMath4.9%50.3Epoch AI
FrontierMath-Tier-4-2025-07-01-PrivateMath4.2%48.0Epoch AI
Critical Physics TasksReasoning0.3%39.0Artificial Analysis

53 benchmarks count, from 63 of 79 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AISierra Research, collected directlyMIT — results are in the licensed repositoryEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be cited

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1