GLM-5

GLM-5 is a non-reasoning model from Z.AI. 34 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.52.2 ±4.3
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.35/s
Input / 1MUS dollars per 1M input tokens.$0.6
Output / 1MUS dollars per 1M output tokens.$1.92
ContextMaximum tokens in one request.205K
EloLMArena rating and rank.1446 (#51)

50 is the middle of the board. The range shows the doubt in the index.

27,605 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
50.7
CodingCode writing and repair.
55.2
ReasoningLogic problems and puzzles.
46.0
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
56.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
56.6
MathMath problems.
52.2

Results

34 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic98.2%69.3Victor Barres et al.
Harvard-MIT Mathematics Tournament February 2025Math97.5%55.7Qwen
Harvard-MIT Mathematics Tournament November 2025Math96.9%Qwen
AIME 2026Math95.8%54.8Qwen
AIME25 first-party comparison snapshotMath93.3%Arcee AI
Instruction-Following EvalInstruction92.6%50.8Jeffrey Zhou et al.
GPQA diamondKnowledge87.8%59.3Epoch AI
τ²-bench TelecomAgentic86.8%61.1enabled effort · Sierra2 Mar 2026Sierra Research
Harvard-MIT Mathematics Tournament February 2026Math86.4%54.8Qwen
Graduate-Level Google-Proof Q&AKnowledge86.0%57.6David Rein et al.
GPQA DiamondKnowledge86.0%57.6David Rein et al.
MMLU-Pro first-party comparison snapshotKnowledge85.8%55.7Arcee AI
Massive Multitask Language Understanding ProfessionalKnowledge85.7%55.6Yubo Wang et al.
MMLU-ProXMultilingual83.1%MMLU-ProX authors
MMAnswerBenchMath82.5%Qwen
τ²-bench AirlineAgentic82.5%58.0enabled effort · Sierra2 Mar 2026Sierra Research
Artificial Analysis GPQA DiamondKnowledge82.0%52.6Artificial Analysis
OTIS Mock AIME 2024-2025Math80.0%56.3Epoch AI
Software Engineering Benchmark VerifiedCoding77.8%59.9Carlos E. Jimenez et al.
Artificial Analysis Long Context ReasoningReasoning75.7%60.6Artificial Analysis
React Native EvalsCoding74.8%52.3Callstack
τ²-bench RetailAgentic73.7%51.7enabled effort · Sierra30 Apr 2026Sierra Research
SWE-bench Verified (mini-swe-agent-v2)Coding72.8%55.9Arcee AI
SWE-bench VerifiedCoding72.8%55.9high effort · mini-SWE-agent1 Sept 2026SWE-bench team
Artificial Analysis IFBenchInstruction72.3%62.5Artificial Analysis
SWE-Bench verifiedCoding72.1%55.4Epoch AI
WideResearchAgentic69.8%54.8Qwen
SWE-bench MultilingualMultilingual69.7%mini-SWE-agent20 Feb 2026SWE-bench team
SWE-bench MultilingualCoding69.7%mini-SWE-agent2 Sept 2026SWE-bench team
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge66.8%52.8Xiaoxuan Du et al.
τ³-Bench Tool-Agent-User EvaluationAgentic65.6%50.9Sierra Research
AI-NeedleReasoning63.3%Qwen
SWE-RebenchCoding62.8%Nebius
MCP-TasksAgentic60.8%Qwen
LongBench v2Reasoning60.8%LongBench v2 authors
Claw-EvalAgentic57.7%49.1Bowen Ye et al.
SWE-bench ProCoding55.1%57.3Xiang Deng et al.
NOVA-63Multilingual55.1%Qwen
QwenClawBenchAgentic54.1%53.4Qwen
Gert Labs Composite Game BenchmarkAgentic51.0%57.8Gert Labs
Humanity's Last ExamKnowledge50.4%71.5Center for AI Safety et al.
ARC-AGI-1 (semi-private)Reasoning44.7%48.3ARC Prize Foundation
CyberGymAgentic43.2%45.8Zhun Wang et al.
ToolathlonAgentic38.0%53.4OpenAI
MCP AtlasAgentic31.1%31.7OpenAI
Artificial Analysis Humanity's Last ExamKnowledge29.3%56.3Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge27.9%57.3Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge26.3%46.4Artificial Analysis
FrontierMath-2025-02-28-PrivateMath16.4%46.0Epoch AI
DeepPlanningAgentic14.6%DeepPlanning authors
APEX-Agents-AAAgentic14.5%52.5Artificial Analysis / Mercor
Chess PuzzlesReasoning10.0%36.5Epoch AI
τ²-bench BankingAgentic9.8%5.9enabled effort · Sierra4 Aug 2026Sierra Research
ARC-AGI-2 (semi-private)Reasoning4.9%42.2ARC Prize Foundation
FrontierMath-Tier-4-2025-07-01-PrivateMath2.1%45.7Epoch AI
Critical Physics TasksReasoning2.0%42.6Artificial Analysis

34 benchmarks count, from 44 of 56 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSierra Research, collected directlyMIT — results are in the licensed repositorySWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be citedARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Z.AI

GLM-5.369.6GLM-5.3-Flash65.6GLM-5.262.7GLM-5.157.8GLM-4.750.2