Z.AI
availableShows if the model has enough results for an index.GLM-5
GLM-5 is a non-reasoning model from Z.AI. 34 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.52.2 ±4.3
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.35/s
Input / 1MUS dollars per 1M input tokens.$0.6
Output / 1MUS dollars per 1M output tokens.$1.92
ContextMaximum tokens in one request.205K
EloLMArena rating and rank.1446 (#51)
50 is the middle of the board. The range shows the doubt in the index.
27,605 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
34 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 98.2% | 69.3 | — | — | Victor Barres et al. |
| Harvard-MIT Mathematics Tournament February 2025 | Math | 97.5% | 55.7 | — | — | Qwen |
| Harvard-MIT Mathematics Tournament November 2025 | Math | 96.9% | — | — | — | Qwen |
| AIME 2026 | Math | 95.8% | 54.8 | — | — | Qwen |
| AIME25 first-party comparison snapshot | Math | 93.3% | — | — | — | Arcee AI |
| Instruction-Following Eval | Instruction | 92.6% | 50.8 | — | — | Jeffrey Zhou et al. |
| GPQA diamond | Knowledge | 87.8% | 59.3 | — | — | Epoch AI |
| τ²-bench Telecom | Agentic | 86.8% | 61.1 | enabled effort · Sierra | 2 Mar 2026 | Sierra Research |
| Harvard-MIT Mathematics Tournament February 2026 | Math | 86.4% | 54.8 | — | — | Qwen |
| Graduate-Level Google-Proof Q&A | Knowledge | 86.0% | 57.6 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 86.0% | 57.6 | — | — | David Rein et al. |
| MMLU-Pro first-party comparison snapshot | Knowledge | 85.8% | 55.7 | — | — | Arcee AI |
| Massive Multitask Language Understanding Professional | Knowledge | 85.7% | 55.6 | — | — | Yubo Wang et al. |
| MMLU-ProX | Multilingual | 83.1% | — | — | — | MMLU-ProX authors |
| MMAnswerBench | Math | 82.5% | — | — | — | Qwen |
| τ²-bench Airline | Agentic | 82.5% | 58.0 | enabled effort · Sierra | 2 Mar 2026 | Sierra Research |
| Artificial Analysis GPQA Diamond | Knowledge | 82.0% | 52.6 | — | — | Artificial Analysis |
| OTIS Mock AIME 2024-2025 | Math | 80.0% | 56.3 | — | — | Epoch AI |
| Software Engineering Benchmark Verified | Coding | 77.8% | 59.9 | — | — | Carlos E. Jimenez et al. |
| Artificial Analysis Long Context Reasoning | Reasoning | 75.7% | 60.6 | — | — | Artificial Analysis |
| React Native Evals | Coding | 74.8% | 52.3 | — | — | Callstack |
| τ²-bench Retail | Agentic | 73.7% | 51.7 | enabled effort · Sierra | 30 Apr 2026 | Sierra Research |
| SWE-bench Verified (mini-swe-agent-v2) | Coding | 72.8% | 55.9 | — | — | Arcee AI |
| SWE-bench Verified | Coding | 72.8% | 55.9 | high effort · mini-SWE-agent | 1 Sept 2026 | SWE-bench team |
| Artificial Analysis IFBench | Instruction | 72.3% | 62.5 | — | — | Artificial Analysis |
| SWE-Bench verified | Coding | 72.1% | 55.4 | — | — | Epoch AI |
| WideResearch | Agentic | 69.8% | 54.8 | — | — | Qwen |
| SWE-bench Multilingual | Multilingual | 69.7% | — | mini-SWE-agent | 20 Feb 2026 | SWE-bench team |
| SWE-bench Multilingual | Coding | 69.7% | — | mini-SWE-agent | 2 Sept 2026 | SWE-bench team |
| SuperGPQA: Scaling LLM Evaluation Across 285 Graduate Disciplines | Knowledge | 66.8% | 52.8 | — | — | Xiaoxuan Du et al. |
| τ³-Bench Tool-Agent-User Evaluation | Agentic | 65.6% | 50.9 | — | — | Sierra Research |
| AI-Needle | Reasoning | 63.3% | — | — | — | Qwen |
| SWE-Rebench | Coding | 62.8% | — | — | — | Nebius |
| MCP-Tasks | Agentic | 60.8% | — | — | — | Qwen |
| LongBench v2 | Reasoning | 60.8% | — | — | — | LongBench v2 authors |
| Claw-Eval | Agentic | 57.7% | 49.1 | — | — | Bowen Ye et al. |
| SWE-bench Pro | Coding | 55.1% | 57.3 | — | — | Xiang Deng et al. |
| NOVA-63 | Multilingual | 55.1% | — | — | — | Qwen |
| QwenClawBench | Agentic | 54.1% | 53.4 | — | — | Qwen |
| Gert Labs Composite Game Benchmark | Agentic | 51.0% | 57.8 | — | — | Gert Labs |
| Humanity's Last Exam | Knowledge | 50.4% | 71.5 | — | — | Center for AI Safety et al. |
| ARC-AGI-1 (semi-private) | Reasoning | 44.7% | 48.3 | — | — | ARC Prize Foundation |
| CyberGym | Agentic | 43.2% | 45.8 | — | — | Zhun Wang et al. |
| Toolathlon | Agentic | 38.0% | 53.4 | — | — | OpenAI |
| MCP Atlas | Agentic | 31.1% | 31.7 | — | — | OpenAI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 29.3% | 56.3 | — | — | Artificial Analysis |
| Artificial Analysis Intelligence Index | Knowledge | 27.9% | 57.3 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 26.3% | 46.4 | — | — | Artificial Analysis |
| FrontierMath-2025-02-28-Private | Math | 16.4% | 46.0 | — | — | Epoch AI |
| DeepPlanning | Agentic | 14.6% | — | — | — | DeepPlanning authors |
| APEX-Agents-AA | Agentic | 14.5% | 52.5 | — | — | Artificial Analysis / Mercor |
| Chess Puzzles | Reasoning | 10.0% | 36.5 | — | — | Epoch AI |
| τ²-bench Banking | Agentic | 9.8% | 5.9 | enabled effort · Sierra | 4 Aug 2026 | Sierra Research |
| ARC-AGI-2 (semi-private) | Reasoning | 4.9% | 42.2 | — | — | ARC Prize Foundation |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 2.1% | 45.7 | — | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 2.0% | 42.6 | — | — | Artificial Analysis |
34 benchmarks count, from 44 of 56 results. A grey row does not count. Too few models took that benchmark.