Anthropic
availableShows if the model has enough results for an index.Claude Haiku 4.5
Claude Haiku 4.5 is a non-reasoning model from Anthropic. 13 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.43.6 ±6.0
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.54/s
Input / 1MUS dollars per 1M input tokens.$1 batch $0.5
Output / 1MUS dollars per 1M output tokens.$5 batch $2.5 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.200K
EloLMArena rating and rank.1397 (#147)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
129,278 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
13 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| MATH level 5 | Math | 91.6% | 43.7 | — | — | Epoch AI |
| VulcanBench v3 | Coding | 76.2% | 55.2 | — | — | VulcanBench contributors |
| Software Engineering Benchmark Verified | Coding | 73.3% | 56.3 | — | — | Carlos E. Jimenez et al. |
| SWE-bench Verified | Coding | 66.6% | 51.0 | high effort · mini-SWE-agent | 1 Sept 2026 | SWE-bench team |
| GPQA diamond | Knowledge | 65.8% | 39.0 | — | — | Epoch AI |
| SWE-bench Multilingual | Coding | 64.7% | — | mini-SWE-agent | 2 Sept 2026 | SWE-bench team |
| EuroEval French | Multilingual | 57.6% | 74.2 | — | — | EuroEval |
| EuroEval Swedish | Multilingual | 53.7% | 69.3 | — | — | EuroEval |
| EuroEval Portuguese | Multilingual | 53.4% | 69.0 | — | — | EuroEval |
| EuroEval Dutch | Multilingual | 53.4% | 68.9 | — | — | EuroEval |
| EuroEval Italian | Multilingual | 53.1% | 68.6 | — | — | EuroEval |
| OTIS Mock AIME 2024-2025 | Math | 51.3% | 40.2 | — | — | Epoch AI |
| EuroEval Spanish | Multilingual | 49.7% | 64.3 | — | — | EuroEval |
| EuroEval Polish | Multilingual | 49.5% | 64.1 | — | — | EuroEval |
| EuroEval German | Multilingual | 47.8% | 62.0 | — | — | EuroEval |
| JobBench | Agentic | 16.0% | 45.1 | — | — | Yuetai Li et al. |
| ARC-AGI-1 (semi-private) | Reasoning | 14.3% | 34.0 | — | — | ARC Prize Foundation |
| SimpleQA Verified | Knowledge | 12.9% | 33.5 | — | — | Epoch AI |
| Chess Puzzles | Reasoning | 8.0% | 33.9 | — | — | Epoch AI |
| FrontierMath-2025-02-28-Private | Math | 5.0% | 35.3 | — | — | Epoch AI |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 2.1% | 45.7 | — | — | Epoch AI |
| ARC-AGI-2 (semi-private) | Reasoning | 1.3% | 40.3 | — | — | ARC Prize Foundation |
13 benchmarks count, from 21 of 22 results. A grey row does not count. Too few models took that benchmark.