Z.AI
availableShows if the model has enough results for an index.GLM-5.2
GLM-5.2 is a reasoning model from Z.AI in the GLM-5 family. 48 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.62.7 ±3.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.74/s
Input / 1MUS dollars per 1M input tokens.$0.65
Output / 1MUS dollars per 1M output tokens.$2.04
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1467 (#29)
50 is the middle of the board. The range shows the doubt in the index. A free tier is available.
36,798 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
48 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| AIME 2026 | Math | 99.2% | 57.3 | — | — | Qwen |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 99.1% | 69.9 | — | — | Victor Barres et al. |
| Harvard-MIT Mathematics Tournament November 2025 | Math | 94.4% | — | — | — | Qwen |
| Harvard-MIT Mathematics Tournament February 2026 | Math | 92.5% | 59.4 | — | — | Qwen |
| GPQA diamond | Knowledge | 91.9% | 63.0 | max effort | — | Epoch AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 91.2% | 62.4 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 91.2% | 62.4 | — | — | David Rein et al. |
| MMAnswerBench | Math | 91.0% | — | — | — | Qwen |
| LiveBench Mathematics | Math | 89.8% | 67.9 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis GPQA Diamond | Knowledge | 89.5% | 60.3 | — | — | Artificial Analysis |
| MMLU Pro | Knowledge | 86.7% | 57.2 | — | 1 Sept 2026 | Vals AI |
| OTIS Mock AIME 2024-2025 | Math | 86.4% | 59.8 | max effort | — | Epoch AI |
| GPQA Diamond | Knowledge | 85.6% | 57.3 | — | 1 Sept 2026 | Vals AI |
| SWE-bench | Coding | 82.8% | 64.0 | — | 1 Sept 2026 | Vals AI |
| LiveBench Coding | Coding | 79.7% | 69.9 | — | 25 Jun 2026 | LiveBench |
| SWE-Bench verified | Coding | 78.7% | 60.7 | max effort | — | Epoch AI |
| LiveBench Reasoning | Reasoning | 78.6% | 65.2 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis Long Context Reasoning | Reasoning | 78.3% | 62.4 | — | — | Artificial Analysis |
| ARC-AGI-1 (semi-private) | Reasoning | 77.0% | 63.6 | — | — | ARC Prize Foundation |
| MCP Atlas | Agentic | 76.8% | 66.0 | — | — | OpenAI |
| LiveBench Language | Knowledge | 76.2% | 62.6 | — | 25 Jun 2026 | LiveBench |
| LiveBench Data Analysis | Reasoning | 73.7% | 58.4 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis IFBench | Instruction | 73.3% | 63.5 | — | — | Artificial Analysis |
| LiveCodeBench | Coding | 69.5% | 47.5 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis Coding Index | Coding | 68.8% | 67.4 | — | — | Artificial Analysis |
| Terminal-Bench 2.1 | Agentic | 67.8% | 64.1 | — | 21 Sept 2026 | Vals AI |
| Vibe Code Bench v1.1 | Coding | 64.0% | 68.8 | OpenHands | 21 Sept 2026 | Vals AI |
| ProgramBench: Can Language Models Rebuild Programs From Scratch? | Coding | 63.7% | 65.5 | — | — | John Yang et al. |
| LiveBench Instruction Following | Instruction | 62.3% | 58.5 | — | 25 Jun 2026 | LiveBench |
| SWE-bench Pro | Coding | 62.1% | 64.1 | — | — | Xiang Deng et al. |
| FrontierMath-Tiers-1-3-v2-Private | Math | 59.2% | 64.8 | max effort | — | Epoch AI |
| OpenHarmony Bench v1.0 | Coding | 58.4% | 67.5 | — | — | OpenHarmony Bench authors |
| cursorBench32 | Coding | 55.0% | 63.8 | — | — | Benchmark authors |
| Humanity's Last Exam | Knowledge | 54.7% | 75.1 | — | — | Center for AI Safety et al. |
| LiveBench Agentic Coding | Agentic | 51.8% | 65.4 | — | 25 Jun 2026 | LiveBench |
| Artificial Analysis SciCode | Coding | 51.2% | 63.9 | — | — | Artificial Analysis |
| NL2Repo | Coding | 48.9% | 63.4 | — | — | MiniMax |
| Toolathlon | Agentic | 48.2% | 62.9 | — | — | OpenAI |
| SkillsBench | Coding | 45.1% | 61.2 | OpenHands | 11 Sept 2026 | Vals AI |
| GDPval-AA normalized | Agentic | 42.9% | 68.5 | — | — | Artificial Analysis |
| Artificial Analysis ITBench-AA | Agentic | 42.7% | — | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 41.1% | 69.1 | — | — | Artificial Analysis |
| Humanity's Last Exam without tools | Knowledge | 40.5% | 63.1 | — | — | OpenAI |
| Artificial Analysis Agentic Index | Agentic | 39.4% | 69.3 | — | — | Artificial Analysis |
| Code Migration | Coding | 37.9% | 69.2 | — | 21 Sept 2026 | Vals AI |
| τ²-bench Banking | Agentic | 37.1% | 25.5 | xhigh effort · Sierra | 4 Aug 2026 | Sierra Research |
| SimpleQA Verified | Knowledge | 34.2% | 53.2 | max effort | — | Epoch AI |
| Artificial Analysis Intelligence Index | Knowledge | 33.7% | 64.6 | — | — | Artificial Analysis |
| APEX-Agents-AA | Agentic | 33.7% | 67.7 | — | — | Artificial Analysis / Mercor |
| FrontierMath-Tier-4-v2-Private | Math | 29.3% | 62.0 | max effort | — | Epoch AI |
| Artificial Analysis Omniscience Accuracy | Knowledge | 24.3% | 43.9 | — | — | Artificial Analysis |
| ARC-AGI-2 (semi-private) | Reasoning | 22.8% | 51.2 | — | — | ARC Prize Foundation |
| Chess Puzzles | Reasoning | 21.0% | 50.7 | max effort | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 20.9% | 82.2 | — | — | Artificial Analysis |
| ResearchClawBench | Agentic | 20.7% | — | — | — | InternScience |
| Mystery Game Puzzles | Reasoning | 15.0% | 49.9 | medium effort | — | Epoch AI |
| EBR-bench | Reasoning | 9.5% | 55.8 | max effort | — | Epoch AI |
| Agent Arena task outcome | Agentic | 4.9 | 72.9 | max effort | 15 Sept 2026 | LMArena |
| Agent Arena steerability | Agentic | 4.9 | 72.9 | max effort | 15 Sept 2026 | LMArena |
| Terminal-Bench 3.0 | Agentic | 4.6% | 58.5 | — | — | Ryan Marten et al. |
| Agent Arena command recovery | Agentic | 1.9 | 69.5 | max effort | 15 Sept 2026 | LMArena |
| ProgramBench | Coding | 0.5% | — | — | 21 Sept 2026 | Vals AI |
48 benchmarks count, from 57 of 62 results. A grey row does not count. Too few models took that benchmark.