Claude Opus 5.5

Claude Opus 5.5 is a reasoning model from Anthropic. 28 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.82.8 ±3.6
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.85/s
Input / 1MUS dollars per 1M input tokens.$4
Output / 1MUS dollars per 1M output tokens.$20
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
80.3
CodingCode writing and repair.
86.3
ReasoningLogic problems and puzzles.
79.2
MultimodalTasks with images and text.
77.1
KnowledgeFacts and expert knowledge.
87.7
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
63.8
MathMath problems.
77.6

Results

28 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
LiveBench MathematicsMath97.1%77.6max effort25 Jun 2026LiveBench
ArXivMath August 2026 with toolsMath96.9%MathArena and Anthropic
BenchCAD Vision2Code voxel IoU with toolsMultimodal96.2%Zhang et al. and Anthropic
Global MMLUMultilingual94.3%Singh et al.
Multi-task Indic Language Understanding BenchmarkMultilingual93.1%Verma et al.
LiveBench ReasoningReasoning92.2%84.0max effort25 Jun 2026LiveBench
Artificial Analysis Harvey LAB-AAAgentic91.2%74.3Artificial Analysis
Legal Agent Benchmark mean criterion-pass rate — Harvey held-out setAgentic91.2%Harvey AI
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding91.2%84.6John Yang et al.
ArXivMath August 2026 without toolsMath91.2%MathArena and Anthropic
SWE-bench ProCoding89.9%91.0Xiang Deng et al.
BioMysteryBench Human SolvableKnowledge89.3%Anthropic
LiveBench CodingCoding89.3%85.8max effort25 Jun 2026LiveBench
Chartography with image and code toolsMultimodal89.0%Surge AI and Anthropic
Artificial Analysis MMMU-ProMultimodal87.7%74.3Artificial Analysis
LiveBench LanguageKnowledge86.3%74.6max effort25 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning84.7%66.8Artificial Analysis
Anthropic de novo protein-binder design evaluationKnowledge82.6%Anthropic
Toolathlon Verified Pass@3Agentic82.4%Anthropic
LiveBench Data AnalysisReasoning80.3%67.5max effort25 Jun 2026LiveBench
OfficeQAMultimodal78.9%Databricks and Anthropic
Toolathlon-VerifiedAgentic77.8%75.4Moonshot AI
HealthBench Professional raw scoreKnowledge77.1%Anthropic
DeepSWEAgentic74.2%77.7Datacurve AI
Molecular Biology Protocols TroubleshootingKnowledge73.7%Anthropic
BenchCAD Vision2Code voxel IoU without toolsMultimodal73.0%Zhang et al. and Anthropic
Toolathlon Verified Pass cubedAgentic72.2%Anthropic
LatchBio SpatialBench VerifiedKnowledge72.0%LatchBio and Anthropic
LiveBench Agentic CodingAgentic71.7%84.2max effort25 Jun 2026LiveBench
Anthropic biomedical-image-analysis evaluationMultimodal71.4%Anthropic
Artificial Analysis AutomationBenchAgentic69.5%84.5Artificial Analysis
Benchling Molecular Biology Protocols UnderstandingKnowledge69.0%Benchling and Anthropic
HealthBench raw scoreKnowledge68.1%Anthropic
Humanity's Last Exam with toolsAgentic67.7%80.1DeepSeek-AI
OfficeQA ProMultimodal67.7%80.0OfficeQA Pro authors
GDPval-AA normalizedAgentic67.3%87.3Artificial Analysis
Artificial Analysis SciCodeCoding66.9%85.7Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge66.2%95.0Artificial Analysis
LiveBench Instruction FollowingInstruction65.7%63.8max effort25 Jun 2026LiveBench
HealthBench ProfessionalKnowledge65.6%Rebecca Soskin Hicks et al.
Chartography without toolsMultimodal64.4%Surge AI and Anthropic
Humanity's Last Exam without toolsKnowledge64.4%83.3OpenAI
FrontierCode 1.1 ExtendedCoding63.6%Cognition
Anthropic medicinal-chemistry evaluationKnowledge63.5%Anthropic
FrontierSWE v2Coding62.3%88.9Proximal
Artificial Analysis Humanity's Last ExamKnowledge61.4%91.1Artificial Analysis
LatchBio SingleCellBenchKnowledge61.2%LatchBio and Anthropic
HealthBench length-adjusted scoreKnowledge60.6%Anthropic
Anthropic Protein Design evaluationKnowledge60.2%Anthropic
Terminal-Bench-Science 0.1Agentic58.7%Terminal-Bench-Science Team
cursorBench40Coding57.8%Benchmark authors
Artificial Analysis Intelligence IndexKnowledge57.6%94.5Artificial Analysis
Anthropic protein-design library-ranking taskKnowledge56.0%Anthropic
FrontierCode 1.1 MainCoding54.4%82.0Cognition
BioMysteryBench Human DifficultKnowledge50.0%Anthropic
OSWorld 2.0Agentic48.7%79.1Mengqi Yuan et al.
AutomationBenchAgentic40.0%82.4Moonshot AI
Axiom Bio morphology-to-molecule matchingKnowledge34.0%Axiom Bio and Anthropic
Critical Physics TasksReasoning31.7%95.0Artificial Analysis
Toolathlon Verified average assistant turnsAgentic26.9%Anthropic
Artificial Analysis GDP.pdfAgentic26.2%77.7Artificial Analysis
Legal Agent Benchmark all-pass rate — Harvey held-out setAgentic8.3%Harvey AI

28 benchmarks count, from 29 of 62 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not results

More from Anthropic

Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1Claude Sonnet 567.6