GLM-5.3-Flash

GLM-5.3-Flash is a reasoning model from Z.AI in the GLM-5 family. 41 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.65.6 ±2.5
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.69/s
Input / 1MUS dollars per 1M input tokens.$0.15
Output / 1MUS dollars per 1M output tokens.$0.5
ContextMaximum tokens in one request.1.31M
EloLMArena rating and rank.1472 (#24)

50 is the middle of the board. The range shows the doubt in the index.

10,038 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
71.8
CodingCode writing and repair.
64.1
ReasoningLogic problems and puzzles.
57.4
MultimodalTasks with images and text.
70.0
KnowledgeFacts and expert knowledge.
63.3
MultilingualTasks in many languages.
91.4
InstructionTasks with strict rules in the prompt.
43.7
MathMath problems.
59.5

Results

41 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math93.9%64.0max effortEpoch AI
SWE-benchCoding92.0%71.31 Sept 2026Vals AI
Artificial Analysis GPQA DiamondKnowledge91.2%62.0Artificial Analysis
GPQA diamondKnowledge90.2%61.5max effortEpoch AI
CharXiv ReasoningMultimodal89.4%67.9CharXiv authors
GPQA DiamondKnowledge86.4%58.01 Sept 2026Vals AI
MMLU ProKnowledge86.1%56.11 Sept 2026Vals AI
MMMU ProMultimodal86.0%67.31 Sept 2026Vals AI
Terminal-Bench 2.1 (provider run)Agentic84.3%73.8DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic84.3%73.8DeepSeek-AI
LiveBench MathematicsMath81.2%56.425 Jun 2026LiveBench
LiveCodeBenchCoding80.5%57.51 Sept 2026Vals AI
Multimodal Multi-disciplinary Video UnderstandingMultimodal80.5%MMVU benchmark maintainers
LiveBench CodingCoding79.0%68.725 Jun 2026LiveBench
Toolathlon-VerifiedAgentic78.4%75.9Moonshot AI
Chartography with image and code toolsMultimodal78.0%Surge AI and Anthropic
LiveBench ReasoningReasoning77.6%63.825 Jun 2026LiveBench
LiveBench LanguageKnowledge77.3%63.825 Jun 2026LiveBench
LiveBench Data AnalysisReasoning76.4%62.125 Jun 2026LiveBench
EuroEval PortugueseMultilingual74.1%94.8EuroEval
EuroEval FrenchMultilingual73.5%94.0EuroEval
EuroEval ItalianMultilingual73.3%93.7EuroEval
EuroEval SwedishMultilingual72.8%93.2EuroEval
EuroEval DutchMultilingual69.9%89.6EuroEval
EuroEval PolishMultilingual68.1%87.3EuroEval
EuroEval SpanishMultilingual67.4%86.5EuroEval
EuroEval GermanMultilingual63.6%81.7EuroEval
DeepSWEAgentic63.4%69.9Datacurve AI
Terminal-Bench 2.1Agentic62.9%61.221 Sept 2026Vals AI
OfficeQA ProMultimodal62.4%74.7OfficeQA Pro authors
Artificial Analysis AutomationBenchAgentic60.4%72.5Artificial Analysis
OpenHarmony Bench v1.0Coding57.3%66.4OpenHarmony Bench authors
LiveBench Agentic CodingAgentic56.8%70.125 Jun 2026LiveBench
NL2RepoCoding56.3%69.3MiniMax
FrontierMath-Tiers-1-3-v2-PrivateMath55.8%62.8max effortEpoch AI
Humanity's Last Exam with toolsAgentic55.3%68.1DeepSeek-AI
BabyVisionMultimodal53.4%Meta AI
LiveBench Instruction FollowingInstruction52.8%43.725 Jun 2026LiveBench
IOICoding52.5%68.521 Sept 2026Vals AI
Medical Long Context Reasoning (MLCR-AA)Reasoning51.1%82.5Wisedocs and Artificial Analysis
AutomationBenchAgentic48.8%95.0Moonshot AI
Artificial Analysis Tau3-BankingAgentic47.2%76.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge41.8%74.7Artificial Analysis
SkillsBenchCoding40.2%57.0OpenHands11 Sept 2026Vals AI
Artificial Analysis EnterpriseOps-GymAgentic33.2%52.9Artificial Analysis
Vibe Code Bench v1.1Coding30.8%55.0OpenHands21 Sept 2026Vals AI
Agents' Last ExamAgentic26.3%66.0DeepSeek-AI
ProofBench v1.1Math21.0%57.921 Sept 2026Vals AI
Code MigrationCoding20.5%58.021 Sept 2026Vals AI
Terminal-Bench 4.0Agentic19.7%73.321 Sept 2026Vals AI
FrontierMath-Tier-4-v2-PrivateMath17.1%56.1max effortEpoch AI
Vibe Code Bench 1-100Coding16.0%69.4OpenHands16 Sept 2026Vals AI
Chess PuzzlesReasoning14.0%41.6max effortEpoch AI
Agent Arena task outcomeAgentic8.276.715 Sept 2026LMArena
Mystery Game PuzzlesReasoning8.0%42.4max effortEpoch AI
Agent Arena steerabilityAgentic0.167.515 Sept 2026LMArena
ProgramBenchCoding0.0%21 Sept 2026Vals AI
Agent Arena command recoveryAgentic-4.662.115 Sept 2026LMArena

41 benchmarks count, from 54 of 58 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from Z.AI

GLM-5.369.6GLM-5.262.7GLM-5.157.8GLM-552.2GLM-4.750.2