Anthropic
availableShows if the model has enough results for an index.Claude Sonnet 4.5
Claude Sonnet 4.5 is a non-reasoning model from Anthropic. 20 benchmarks count toward its score, in 5 categories.
IndexOverall score. 50 is the middle.47.7 ±5.6
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.39/s
Input / 1MUS dollars per 1M input tokens.$3 batch $1.5
Output / 1MUS dollars per 1M output tokens.$15 batch $7.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1438 (#72)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
79,628 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
20 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| MATH level 5 | Math | 97.7% | 46.7 | — | — | Epoch AI |
| τ²-bench Telecom | Agentic | 91.4% | 64.4 | enabled effort · Sierra | 2 Mar 2026 | Sierra Research |
| American Invitational Mathematics Examination 2025 | Math | 87.0% | 45.3 | — | — | Mathematical Association of America |
| Graduate-Level Google-Proof Q&A | Knowledge | 83.4% | 55.2 | — | — | David Rein et al. |
| GPQA diamond | Knowledge | 80.3% | 52.3 | — | — | Epoch AI |
| τ²-bench Retail | Agentic | 79.3% | 55.7 | enabled effort · Sierra | 30 Apr 2026 | Sierra Research |
| Software Engineering Benchmark Verified | Coding | 77.2% | 59.5 | — | — | Carlos E. Jimenez et al. |
| OTIS Mock AIME 2024-2025 | Math | 74.4% | 53.2 | — | — | Epoch AI |
| SWE-bench Verified | Coding | 71.4% | 54.8 | high effort · mini-SWE-agent | 1 Sept 2026 | SWE-bench team |
| SWE-Bench verified | Coding | 71.3% | 54.7 | — | — | Epoch AI |
| τ²-bench Airline | Agentic | 71.0% | 49.8 | enabled effort · Sierra | 2 Mar 2026 | Sierra Research |
| SWE-bench Multilingual | Coding | 67.0% | — | mini-SWE-agent | 2 Sept 2026 | SWE-bench team |
| OSWorld-Verified | Agentic | 61.4% | 50.9 | — | — | Tianbao Xie et al. |
| Gert Labs Composite Game Benchmark | Agentic | 48.5% | 55.6 | — | — | Gert Labs |
| JobBench | Agentic | 27.7% | 53.2 | — | — | Yuetai Li et al. |
| SimpleQA Verified | Knowledge | 27.2% | 46.7 | — | — | Epoch AI |
| ARC-AGI-1 (semi-private) | Reasoning | 25.5% | 39.3 | — | — | ARC Prize Foundation |
| τ²-bench Banking | Agentic | 25.3% | 17.0 | enabled effort · Sierra | 4 Aug 2026 | Sierra Research |
| FrontierMath-Tiers-1-3-v2-Private | Math | 23.9% | 44.9 | — | — | Epoch AI |
| VITA-Bench | Agentic | 17.0% | 37.0 | — | — | Meituan LongCat Team |
| Mystery Game Puzzles | Reasoning | 17.0% | 52.0 | — | — | Epoch AI |
| FrontierMath-2025-02-28-Private | Math | 13.5% | 43.2 | — | — | Epoch AI |
| Chess Puzzles | Reasoning | 8.0% | 33.9 | — | — | Epoch AI |
| ARC-AGI-2 (semi-private) | Reasoning | 3.8% | 41.6 | — | — | ARC Prize Foundation |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 3.1% | 46.8 | — | — | Epoch AI |
| FrontierMath-Tier-4-v2-Private | Math | 2.4% | 49.1 | — | — | Epoch AI |
| EBR-bench | Reasoning | 2.4% | 51.0 | — | — | Epoch AI |
20 benchmarks count, from 26 of 27 results. A grey row does not count. Too few models took that benchmark.