Claude Sonnet 4.5

Claude Sonnet 4.5 is a non-reasoning model from Anthropic. 20 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.47.7 ±5.6
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.39/s
Input / 1MUS dollars per 1M input tokens.$3 batch $1.5
Output / 1MUS dollars per 1M output tokens.$15 batch $7.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1438 (#72)

50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.

79,628 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
49.9
CodingCode writing and repair.
54.8
ReasoningLogic problems and puzzles.
43.5
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
50.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
47.0

Results

20 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MATH level 5Math97.7%46.7Epoch AI
τ²-bench TelecomAgentic91.4%64.4enabled effort · Sierra2 Mar 2026Sierra Research
American Invitational Mathematics Examination 2025Math87.0%45.3Mathematical Association of America
Graduate-Level Google-Proof Q&AKnowledge83.4%55.2David Rein et al.
GPQA diamondKnowledge80.3%52.3Epoch AI
τ²-bench RetailAgentic79.3%55.7enabled effort · Sierra30 Apr 2026Sierra Research
Software Engineering Benchmark VerifiedCoding77.2%59.5Carlos E. Jimenez et al.
OTIS Mock AIME 2024-2025Math74.4%53.2Epoch AI
SWE-bench VerifiedCoding71.4%54.8high effort · mini-SWE-agent1 Sept 2026SWE-bench team
SWE-Bench verifiedCoding71.3%54.7Epoch AI
τ²-bench AirlineAgentic71.0%49.8enabled effort · Sierra2 Mar 2026Sierra Research
SWE-bench MultilingualCoding67.0%mini-SWE-agent2 Sept 2026SWE-bench team
OSWorld-VerifiedAgentic61.4%50.9Tianbao Xie et al.
Gert Labs Composite Game BenchmarkAgentic48.5%55.6Gert Labs
JobBenchAgentic27.7%53.2Yuetai Li et al.
SimpleQA VerifiedKnowledge27.2%46.7Epoch AI
ARC-AGI-1 (semi-private)Reasoning25.5%39.3ARC Prize Foundation
τ²-bench BankingAgentic25.3%17.0enabled effort · Sierra4 Aug 2026Sierra Research
FrontierMath-Tiers-1-3-v2-PrivateMath23.9%44.9Epoch AI
VITA-BenchAgentic17.0%37.0Meituan LongCat Team
Mystery Game PuzzlesReasoning17.0%52.0Epoch AI
FrontierMath-2025-02-28-PrivateMath13.5%43.2Epoch AI
Chess PuzzlesReasoning8.0%33.9Epoch AI
ARC-AGI-2 (semi-private)Reasoning3.8%41.6ARC Prize Foundation
FrontierMath-Tier-4-2025-07-01-PrivateMath3.1%46.8Epoch AI
FrontierMath-Tier-4-v2-PrivateMath2.4%49.1Epoch AI
EBR-benchReasoning2.4%51.0Epoch AI

20 benchmarks count, from 26 of 27 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSierra Research, collected directlyMIT — results are in the licensed repositorySWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be citedARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Anthropic

Claude Opus 5.582.8Claude Fable 5.181.4Claude Opus 577.9Claude Fable 577.5Claude Opus 4.871.1