Anthropic
availableShows if the model has enough results for an index.Claude Fable 5.1
Claude Fable 5.1 is a reasoning model from Anthropic in the Claude Fable family. 55 benchmarks count toward its score, in 7 categories.
IndexOverall score. 50 is the middle.81.4 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.49/s
Input / 1MUS dollars per 1M input tokens.$10 batch $5
Output / 1MUS dollars per 1M output tokens.$50 batch $25 US dollars per 1M output tokens in a batch.
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1508 (#1)
50 is the middle of the board. The range shows the doubt in the index. Batch work costs less.
5,783 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
55 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| ProofBench v1.1 | Math | 100.0% | 89.9 | — | 21 Sept 2026 | Vals AI |
| OTIS Mock AIME 2024-2025 | Math | 100.0% | 67.4 | max effort | — | Epoch AI |
| ARC-AGI-1 (semi-private) | Reasoning | 97.5% | 73.2 | max effort | — | ARC Prize Foundation |
| LiveBench Mathematics | Math | 97.0% | 77.5 | max effort | 25 Jun 2026 | LiveBench |
| Artificial Analysis GPQA Diamond | Knowledge | 93.7% | 64.6 | — | — | Artificial Analysis |
| GPQA Diamond | Knowledge | 93.4% | 64.5 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Harvey LAB-AA | Agentic | 93.0% | 76.8 | — | — | Artificial Analysis |
| MMLU Pro | Knowledge | 92.4% | 66.1 | — | 1 Sept 2026 | Vals AI |
| LiveBench Reasoning | Reasoning | 91.7% | 83.3 | max effort | 25 Jun 2026 | LiveBench |
| IOI | Coding | 90.8% | 85.2 | — | 21 Sept 2026 | Vals AI |
| MMMU Pro | Multimodal | 90.6% | 74.8 | — | 1 Sept 2026 | Vals AI |
| LiveCodeBench | Coding | 90.5% | 66.6 | — | 1 Sept 2026 | Vals AI |
| Vibe Code Bench v1.1 | Coding | 90.3% | 79.7 | OpenHands | 21 Sept 2026 | Vals AI |
| FrontierMath-Tiers-1-3-v2-Private | Math | 90.2% | 82.2 | max effort | — | Epoch AI |
| ARC-AGI-2 (semi-private) | Reasoning | 90.0% | 85.1 | max effort | — | ARC Prize Foundation |
| LiveBench Language | Knowledge | 89.5% | 78.5 | max effort | 25 Jun 2026 | LiveBench |
| FrontierMath-Tier-4-v2-Private | Math | 87.8% | 90.1 | max effort | — | Epoch AI |
| ProgramBench: Can Language Models Rebuild Programs From Scratch? | Coding | 87.6% | 82.1 | — | — | John Yang et al. |
| LiveBench Coding | Coding | 86.4% | 81.0 | max effort | 25 Jun 2026 | LiveBench |
| Artificial Analysis Long Context Reasoning | Reasoning | 85.3% | 67.3 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 85.0% | 74.2 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 81.6% | 76.5 | — | — | Artificial Analysis |
| Toolathlon Verified Pass@3 | Agentic | 81.5% | — | — | — | Anthropic |
| SWE-bench Pro | Coding | 81.2% | 82.6 | — | — | Xiang Deng et al. |
| LiveBench Data Analysis | Reasoning | 80.3% | 67.5 | max effort | 25 Jun 2026 | LiveBench |
| Toolathlon-Verified | Agentic | 77.8% | 75.4 | — | — | Moonshot AI |
| cursorBench32 | Coding | 73.4% | 81.1 | — | — | Benchmark authors |
| MirrorCode | Coding | 73.3% | — | high effort | — | Epoch AI |
| Toolathlon Verified Pass cubed | Agentic | 73.1% | — | — | — | Anthropic |
| LiveBench Instruction Following | Instruction | 73.0% | 75.2 | max effort | 25 Jun 2026 | LiveBench |
| ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable job | Agentic | 72.0% | 95.0 | — | — | NeoCognition |
| Medical Long Context Reasoning (MLCR-AA) | Reasoning | 71.1% | 95.0 | — | — | Wisedocs and Artificial Analysis |
| SimpleQA Verified | Knowledge | 70.8% | 87.1 | max effort | — | Epoch AI |
| Furniture Assembly | Reasoning | 70.0% | 89.4 | max effort | — | Epoch AI |
| DeepSWE | Agentic | 67.4% | 72.8 | — | — | Datacurve AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 67.2% | 95.0 | — | — | Artificial Analysis |
| LiveBench Agentic Coding | Agentic | 66.1% | 78.9 | max effort | 25 Jun 2026 | LiveBench |
| Humanity's Last Exam | Knowledge | 65.0% | 83.8 | — | — | Center for AI Safety et al. |
| Artificial Analysis SciCode | Coding | 63.1% | 80.4 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 61.7% | 83.0 | — | — | Artificial Analysis |
| SkillsBench | Coding | 61.6% | 75.2 | OpenHands | 11 Sept 2026 | Vals AI |
| Humanity's Last Exam without tools | Knowledge | 60.9% | 80.4 | — | — | OpenAI |
| Artificial Analysis AutomationBench | Agentic | 59.4% | 71.2 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 59.1% | 88.6 | — | — | Artificial Analysis |
| Mystery Game Puzzles | Reasoning | 58.0% | 95.0 | max effort | — | Epoch AI |
| Artificial Analysis Agentic Index | Agentic | 58.0% | 84.4 | — | — | Artificial Analysis |
| Artificial Analysis AnalystAgent | Agentic | 57.5% | 81.9 | — | — | Artificial Analysis |
| EBR-bench | Reasoning | 57.1% | 88.4 | max effort | — | Epoch AI |
| FrontierSWE v2 | Coding | 56.3% | 85.6 | — | — | Proximal |
| Code Migration | Coding | 54.6% | 79.9 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Intelligence Index | Knowledge | 53.4% | 89.1 | — | — | Artificial Analysis |
| Terminal-Bench-Science 0.1 | Agentic | 52.6% | — | — | — | Terminal-Bench-Science Team |
| cursorBench40 | Coding | 51.8% | — | — | — | Benchmark authors |
| Terminal-Bench 4.0 | Agentic | 49.5% | 94.3 | — | 21 Sept 2026 | Vals AI |
| Artificial Analysis Tau3-Banking | Agentic | 47.2% | 76.7 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 47.0% | 84.3 | max effort | — | Epoch AI |
| Bug Hunt Bench | Coding | 43.0% | — | — | — | Pawel Huryn |
| OSWorld 2.0 | Agentic | 41.7% | 75.8 | — | — | Mengqi Yuan et al. |
| AutomationBench | Agentic | 31.4% | 67.7 | — | — | Moonshot AI |
| Critical Physics Tasks | Reasoning | 29.7% | 95.0 | — | — | Artificial Analysis |
| Vibe Code Bench 1-100 | Coding | 28.0% | 82.6 | OpenHands | 16 Sept 2026 | Vals AI |
| Artificial Analysis GDP.pdf | Agentic | 26.2% | 77.7 | — | — | Artificial Analysis |
| Toolathlon Verified average assistant turns | Agentic | 23.7% | — | — | — | Anthropic |
| Agent Arena task outcome | Agentic | 19.8 | 89.7 | max effort | 15 Sept 2026 | LMArena |
| Agent Arena command recovery | Agentic | 12.6 | 81.6 | max effort | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 7.0% | — | — | 21 Sept 2026 | Vals AI |
| Agent Arena steerability | Agentic | 3.9 | 71.7 | max effort | 15 Sept 2026 | LMArena |
| FrontierMath-Erdos | Math | 0.0% | — | max effort | — | Epoch AI |
55 benchmarks count, from 59 of 68 results. A grey row does not count. Too few models took that benchmark.