Alibaba
availableShows if the model has enough results for an index.Qwen3.7 Max
Qwen3.7 Max is a reasoning model from Alibaba. 51 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.62.6 ±3.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.61/s
Input / 1MUS dollars per 1M input tokens.$1.48
Output / 1MUS dollars per 1M output tokens.$4.43
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A
50 is the middle of the board. The range shows the doubt in the index.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
51 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| Harvard-MIT Mathematics Tournament February 2026 | Math | 97.1% | 62.9 | — | — | Qwen |
| OTIS Mock AIME 2024-2025 | Math | 95.6% | 65.0 | — | — | Epoch AI |
| MMLU-Redux | Knowledge | 95.0% | 54.0 | — | — | Qwen |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 94.7% | 66.8 | — | — | Victor Barres et al. |
| Instruction-Following Eval | Instruction | 94.3% | 54.4 | — | — | Jeffrey Zhou et al. |
| Graduate-Level Google-Proof Q&A | Knowledge | 92.4% | 63.5 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 92.4% | 63.5 | — | — | David Rein et al. |
| Artificial Analysis GPQA Diamond | Knowledge | 92.3% | 63.1 | — | — | Artificial Analysis |
| LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code | Coding | 91.6% | 67.6 | — | — | Naman Jain et al. |
| GPQA diamond | Knowledge | 90.9% | 62.2 | — | — | Epoch AI |
| MRCRv2 | Reasoning | 90.4% | — | — | — | OpenAI |
| MMMLU | Knowledge | 90.3% | — | — | — | OpenAI |
| GPQA Diamond | Knowledge | 90.2% | 61.5 | — | 1 Sept 2026 | Vals AI |
| IMOAnswerBench | Math | 90.0% | — | — | — | DeepSeek-AI |
| Massive Multitask Language Understanding Professional | Knowledge | 89.6% | 61.7 | — | — | Yubo Wang et al. |
| MMLU Pro | Knowledge | 89.3% | 61.3 | — | 1 Sept 2026 | Vals AI |
| MAXIFE | Multilingual | 89.2% | — | — | — | Qwen |
| LiveCodeBench | Coding | 87.1% | 63.5 | — | 1 Sept 2026 | Vals AI |
| MMLU-ProX | Multilingual | 87.0% | — | — | — | MMLU-ProX authors |
| PolyMath | Multilingual | 86.5% | — | — | — | Qwen |
| INCLUDE | Multilingual | 86.2% | — | — | — | Qwen |
| LiveBench Mathematics | Math | 85.2% | 61.8 | — | 25 Jun 2026 | LiveBench |
| LiveBench Reasoning | Reasoning | 83.3% | 71.7 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis IFBench | Instruction | 80.5% | 70.8 | — | — | Artificial Analysis |
| Software Engineering Benchmark Verified | Coding | 80.4% | 62.0 | — | — | Carlos E. Jimenez et al. |
| LiveBench Language | Knowledge | 79.7% | 66.8 | — | 25 Jun 2026 | LiveBench |
| Instruction Following Benchmark | Instruction | 79.1% | 58.1 | — | — | Benchmark authors |
| Artificial Analysis Long Context Reasoning | Reasoning | 79.0% | 62.9 | — | — | Artificial Analysis |
| SWE-Bench verified | Coding | 77.3% | 59.5 | — | — | Epoch AI |
| MCP Atlas | Agentic | 76.4% | 65.7 | — | — | OpenAI |
| Berkeley Function Calling Leaderboard v4 | Agentic | 75.0% | 60.0 | — | — | Arcee AI |
| LiveBench Coding | Coding | 74.2% | 60.9 | — | 25 Jun 2026 | LiveBench |
| LiveBench Instruction Following | Instruction | 74.0% | 76.8 | — | 25 Jun 2026 | LiveBench |
| SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines | Knowledge | 73.6% | 58.4 | — | — | Xiaoxuan Du et al. |
| LiveBench Data Analysis | Reasoning | 71.8% | 55.7 | — | 25 Jun 2026 | LiveBench |
| SWE-bench | Coding | 68.8% | 52.7 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 66.0% | 65.4 | — | — | Artificial Analysis |
| Claw-Eval | Agentic | 65.2% | 60.9 | — | — | Bowen Ye et al. |
| FrontierMath-Tiers-1-3-v2-Private | Math | 64.6% | 67.8 | — | — | Epoch AI |
| QwenClawBench | Agentic | 64.3% | 62.7 | — | — | Qwen |
| Gert Labs Composite Game Benchmark | Agentic | 64.3% | 69.5 | — | — | Gert Labs |
| Terminal-Bench 2.1 | Agentic | 61.0% | 60.1 | — | 21 Sept 2026 | Vals AI |
| SWE-bench Pro | Coding | 60.6% | 62.6 | — | — | Xiang Deng et al. |
| Terminal-Bench 2.0 | Agentic | 59.2% | 64.9 | — | 4 Jun 2026 | Vals AI |
| NOVA-63 | Multilingual | 59.0% | — | — | — | Qwen |
| SimpleQA Verified | Knowledge | 55.8% | 73.2 | — | — | Epoch AI |
| Humanity's Last Exam with tools | Agentic | 53.5% | 66.4 | — | — | DeepSeek-AI |
| Scientific Code Benchmark | Coding | 53.5% | 64.3 | — | — | Benchmark authors |
| OpenHarmony Bench v1.0 | Coding | 53.4% | 62.3 | — | — | OpenHarmony Bench authors |
| Artificial Analysis SciCode | Coding | 49.5% | 61.6 | — | — | Artificial Analysis |
| VITA-Bench | Agentic | 47.9% | 62.7 | — | — | Meituan LongCat Team |
| Vibe Code Bench v1.1 | Coding | 47.7% | 62.0 | OpenHands | 21 Sept 2026 | Vals AI |
| NL2Repo | Coding | 47.2% | 62.0 | — | — | MiniMax |
| IOI v1 | Coding | 46.8% | 66.5 | — | 9 Aug 2026 | Vals AI |
| Apex | Math | 44.5% | — | — | — | DeepSeek-AI |
| LiveBench Agentic Coding | Agentic | 43.6% | 57.7 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis ITBench-AA | Agentic | 42.5% | — | — | — | Artificial Analysis |
| Humanity's Last Exam | Knowledge | 41.4% | 63.9 | — | — | Center for AI Safety et al. |
| Artificial Analysis Humanity's Last Exam | Knowledge | 40.5% | 68.5 | — | — | Artificial Analysis |
| FrontierMath-Tier-4-v2-Private | Math | 34.1% | 64.3 | — | — | Epoch AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 31.1% | 52.3 | — | — | Artificial Analysis |
| GDPval-AA normalized | Agentic | 30.7% | 59.1 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 29.5% | 59.2 | — | — | Artificial Analysis |
| Mystery Game Puzzles | Reasoning | 26.0% | 61.6 | none effort | — | Epoch AI |
| Artificial Analysis Agentic Index | Agentic | 23.9% | 56.7 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 19.0% | 48.1 | — | — | Epoch AI |
| ResearchClawBench | Agentic | 18.7% | — | — | — | InternScience |
| Critical Physics Tasks | Reasoning | 13.4% | 66.5 | — | — | Artificial Analysis |
| Code Migration | Coding | 13.1% | 53.3 | — | 21 Sept 2026 | Vals AI |
| EBR-bench | Reasoning | 9.5% | 55.8 | — | — | Epoch AI |
51 benchmarks count, from 59 of 70 results. A grey row does not count. Too few models took that benchmark.