GLM-5.2

GLM-5.2 is a reasoning model from Z.AI in the GLM-5 family. 48 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.62.7 ±3.8
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.74/s
Input / 1MUS dollars per 1M input tokens.$0.65
Output / 1MUS dollars per 1M output tokens.$2.04
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.1467 (#29)

50 is the middle of the board. The range shows the doubt in the index. A free tier is available.

36,798 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
64.3
CodingCode writing and repair.
64.2
ReasoningLogic problems and puzzles.
59.7
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
60.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
61.0
MathMath problems.
61.9

Results

48 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIME 2026Math99.2%57.3Qwen
τ²-Bench Tool-Agent-User EvaluationAgentic99.1%69.9Victor Barres et al.
Harvard-MIT Mathematics Tournament November 2025Math94.4%Qwen
Harvard-MIT Mathematics Tournament February 2026Math92.5%59.4Qwen
GPQA diamondKnowledge91.9%63.0max effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge91.2%62.4David Rein et al.
GPQA DiamondKnowledge91.2%62.4David Rein et al.
MMAnswerBenchMath91.0%Qwen
LiveBench MathematicsMath89.8%67.925 Jun 2026LiveBench
Artificial Analysis GPQA DiamondKnowledge89.5%60.3Artificial Analysis
MMLU ProKnowledge86.7%57.21 Sept 2026Vals AI
OTIS Mock AIME 2024-2025Math86.4%59.8max effortEpoch AI
GPQA DiamondKnowledge85.6%57.31 Sept 2026Vals AI
SWE-benchCoding82.8%64.01 Sept 2026Vals AI
LiveBench CodingCoding79.7%69.925 Jun 2026LiveBench
SWE-Bench verifiedCoding78.7%60.7max effortEpoch AI
LiveBench ReasoningReasoning78.6%65.225 Jun 2026LiveBench
Artificial Analysis Long Context ReasoningReasoning78.3%62.4Artificial Analysis
ARC-AGI-1 (semi-private)Reasoning77.0%63.6ARC Prize Foundation
MCP AtlasAgentic76.8%66.0OpenAI
LiveBench LanguageKnowledge76.2%62.625 Jun 2026LiveBench
LiveBench Data AnalysisReasoning73.7%58.425 Jun 2026LiveBench
Artificial Analysis IFBenchInstruction73.3%63.5Artificial Analysis
LiveCodeBenchCoding69.5%47.51 Sept 2026Vals AI
Artificial Analysis Coding IndexCoding68.8%67.4Artificial Analysis
Terminal-Bench 2.1Agentic67.8%64.121 Sept 2026Vals AI
Vibe Code Bench v1.1Coding64.0%68.8OpenHands21 Sept 2026Vals AI
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding63.7%65.5John Yang et al.
LiveBench Instruction FollowingInstruction62.3%58.525 Jun 2026LiveBench
SWE-bench ProCoding62.1%64.1Xiang Deng et al.
FrontierMath-Tiers-1-3-v2-PrivateMath59.2%64.8max effortEpoch AI
OpenHarmony Bench v1.0Coding58.4%67.5OpenHarmony Bench authors
cursorBench32Coding55.0%63.8Benchmark authors
Humanity's Last ExamKnowledge54.7%75.1Center for AI Safety et al.
LiveBench Agentic CodingAgentic51.8%65.425 Jun 2026LiveBench
Artificial Analysis SciCodeCoding51.2%63.9Artificial Analysis
NL2RepoCoding48.9%63.4MiniMax
ToolathlonAgentic48.2%62.9OpenAI
SkillsBenchCoding45.1%61.2OpenHands11 Sept 2026Vals AI
GDPval-AA normalizedAgentic42.9%68.5Artificial Analysis
Artificial Analysis ITBench-AAAgentic42.7%Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge41.1%69.1Artificial Analysis
Humanity's Last Exam without toolsKnowledge40.5%63.1OpenAI
Artificial Analysis Agentic IndexAgentic39.4%69.3Artificial Analysis
Code MigrationCoding37.9%69.221 Sept 2026Vals AI
τ²-bench BankingAgentic37.1%25.5xhigh effort · Sierra4 Aug 2026Sierra Research
SimpleQA VerifiedKnowledge34.2%53.2max effortEpoch AI
Artificial Analysis Intelligence IndexKnowledge33.7%64.6Artificial Analysis
APEX-Agents-AAAgentic33.7%67.7Artificial Analysis / Mercor
FrontierMath-Tier-4-v2-PrivateMath29.3%62.0max effortEpoch AI
Artificial Analysis Omniscience AccuracyKnowledge24.3%43.9Artificial Analysis
ARC-AGI-2 (semi-private)Reasoning22.8%51.2ARC Prize Foundation
Chess PuzzlesReasoning21.0%50.7max effortEpoch AI
Critical Physics TasksReasoning20.9%82.2Artificial Analysis
ResearchClawBenchAgentic20.7%InternScience
Mystery Game PuzzlesReasoning15.0%49.9medium effortEpoch AI
EBR-benchReasoning9.5%55.8max effortEpoch AI
Agent Arena task outcomeAgentic4.972.9max effort15 Sept 2026LMArena
Agent Arena steerabilityAgentic4.972.9max effort15 Sept 2026LMArena
Terminal-Bench 3.0Agentic4.6%58.5Ryan Marten et al.
Agent Arena command recoveryAgentic1.969.5max effort15 Sept 2026LMArena
ProgramBenchCoding0.5%21 Sept 2026Vals AI

48 benchmarks count, from 57 of 62 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial useSierra Research, collected directlyMIT — results are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from Z.AI

GLM-5.369.6GLM-5.3-Flash65.6GLM-5.157.8GLM-552.2GLM-4.750.2