Claude Opus 4.7 (Adaptive)

Claude Opus 4.7 (Adaptive) is a reasoning model from Anthropic in the Claude Opus 4.7 family. 24 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.66.2 ±4.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.$5
Output / 1MUS dollars per 1M output tokens.$25
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
66.1
CodingCode writing and repair.
68.3
ReasoningLogic problems and puzzles.
63.1
MultimodalTasks with images and text.
63.0
KnowledgeFacts and expert knowledge.
69.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
48.6
MathMath problems.
N/A

Results

24 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Graduate-Level Google-Proof Q&AKnowledge94.2%65.2David Rein et al.
GPQA DiamondKnowledge94.2%65.2David Rein et al.
Artificial Analysis GPQA DiamondKnowledge91.4%62.2Artificial Analysis
CharXiv ReasoningMultimodal91.0%69.7CharXiv authors
τ²-Bench Tool-Agent-User EvaluationAgentic88.6%62.4Victor Barres et al.
Software Engineering Benchmark VerifiedCoding87.6%67.8Carlos E. Jimenez et al.
CharXiv Reasoning without toolsMultimodal82.1%CharXiv authors
BrowseCompAgentic79.3%66.4OpenAI
Artificial Analysis MMMU-ProMultimodal78.8%63.4Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning78.7%62.7Artificial Analysis
OSWorld-VerifiedAgentic78.0%66.4Tianbao Xie et al.
MCP AtlasAgentic77.3%66.4OpenAI
Artificial Analysis Coding IndexCoding73.6%70.8Artificial Analysis
CyberGymAgentic73.1%66.2Zhun Wang et al.
SWE-bench ProCoding64.3%66.2Xiang Deng et al.
OpenAI MRCR v2 8-needle 128K-256KReasoning59.2%OpenAI
Artificial Analysis IFBenchInstruction58.6%48.6Artificial Analysis
Humanity's Last ExamKnowledge54.7%75.1Center for AI Safety et al.
Artificial Analysis Omniscience AccuracyKnowledge48.9%74.4Artificial Analysis
Humanity's Last Exam without toolsKnowledge46.9%68.5OpenAI
Artificial Analysis ITBench-AAAgentic46.7%Artificial Analysis
JobBenchAgentic45.9%65.7Yuetai Li et al.
OfficeQA ProMultimodal43.6%55.9OfficeQA Pro authors
Artificial Analysis Humanity's Last ExamKnowledge42.3%70.4Artificial Analysis
GDPval-AA normalizedAgentic41.9%67.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge40.7%73.3Artificial Analysis
Artificial Analysis Agentic IndexAgentic39.5%69.4Artificial Analysis
OSWorld 2.0Agentic18.2%64.6Mengqi Yuan et al.
Critical Physics TasksReasoning12.0%63.5Artificial Analysis

24 benchmarks count, from 26 of 29 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 073164.6 · $0.28

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1