Qwen3.8 Max

Qwen3.8 Max is a reasoning model from Alibaba. 48 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.68.4 ±2.8
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.39/s
Input / 1MUS dollars per 1M input tokens.N/A
Output / 1MUS dollars per 1M output tokens.N/A
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
69.4
CodingCode writing and repair.
66.8
ReasoningLogic problems and puzzles.
69.0
MultimodalTasks with images and text.
67.9
KnowledgeFacts and expert knowledge.
64.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
69.6
MathMath problems.
70.7

Results

48 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
OTIS Mock AIME 2024-2025Math99.4%67.1xhigh effortEpoch AI
MathVision with PythonMultimodal97.7%Moonshot AI / MathVision authors
MathVisionMultimodal95.2%Qwen
GPQA DiamondKnowledge93.7%64.71 Sept 2026Vals AI
CharXiv ReasoningMultimodal93.5%72.6CharXiv authors
PaperBenchCoding93.0%Qwen Team
MRCRv2Reasoning92.9%OpenAI
GPQA diamondKnowledge92.7%63.8xhigh effortEpoch AI
Graduate-Level Google-Proof Q&AKnowledge92.6%63.7David Rein et al.
GPQA DiamondKnowledge92.6%63.7David Rein et al.
OmniDocBench 1.5Multimodal92.1%OpenAI
LiveBench MathematicsMath91.3%69.925 Jun 2026LiveBench
BabyVision with PythonMultimodal91.3%Moonshot AI
MLVU mean averageMultimodal90.8%Qwen
Video-MME with subtitleMultimodal90.4%Qwen
VideoMMMUMultimodal88.7%Qwen
MMLU ProKnowledge88.6%60.11 Sept 2026Vals AI
CharXiv Reasoning without toolsMultimodal88.4%CharXiv authors
LiveBench ReasoningReasoning88.2%78.525 Jun 2026LiveBench
MMMU ProMultimodal88.0%70.61 Sept 2026Vals AI
RealWorldQAMultimodal88.0%62.2Qwen
LiveCodeBenchCoding87.9%64.21 Sept 2026Vals AI
Terminal-Bench 2.1 (provider run)Agentic86.6%75.2DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic86.6%75.2DeepSeek-AI
OSWorld-VerifiedAgentic86.1%73.9Tianbao Xie et al.
SWE-benchCoding85.6%66.21 Sept 2026Vals AI
AndroidWorldAgentic85.3%Z.AI
ScreenSpot ProMultimodal84.5%66.5Kaixin Li et al.
Instruction Following BenchmarkInstruction82.8%62.4Benchmark authors
Multimodal Multi-disciplinary Video UnderstandingMultimodal82.4%MMVU benchmark maintainers
Massive Multi-discipline Multimodal Understanding ProMultimodal82.3%61.3MMMU-Pro authors
BabyVisionMultimodal82.0%Meta AI
WideResearchAgentic81.9%67.9Qwen
LVBenchMultimodal81.8%Qwen Team
VulcanBench v3Coding81.2%63.3VulcanBench contributors
MedXpertQA MultimodalMultimodal80.4%Meta AI
LiveBench LanguageKnowledge79.7%66.725 Jun 2026LiveBench
CC-OCRMultimodal79.6%Qwen
LiveBench Data AnalysisReasoning78.4%64.925 Jun 2026LiveBench
MobileWorldAgentic77.8%Qwen Team
ERQAMultimodal77.8%71.7Qwen
SimpleVQAMultimodal75.0%68.2Z.AI
CoWorkBenchAgentic74.8%Qwen Team
FrontierMath-Tiers-1-3-v2-PrivateMath74.7%73.5xhigh effortEpoch AI
OCRBench V2Multimodal74.2%OCRBench authors
LiveBench Instruction FollowingInstruction74.1%76.925 Jun 2026LiveBench
FrontierSWECoding73.5%Evan Chu et al.
IOI v1Coding73.0%81.59 Aug 2026Vals AI
LiveBench CodingCoding72.9%58.725 Jun 2026LiveBench
Toolathlon-VerifiedAgentic72.5%71.0Moonshot AI
Vision2WebMultimodal69.0%Z.AI
IOICoding68.9%75.721 Sept 2026Vals AI
SWE-bench ProCoding67.7%69.5Xiang Deng et al.
Terminal-Bench 2.1Agentic67.4%63.821 Sept 2026Vals AI
WebArena-Verified Browser Agent BenchmarkAgentic66.8%Amine El Hattami et al.
LongBench v2Reasoning66.3%LongBench v2 authors
Vibe Code Bench v1.1Coding64.7%69.1OpenHands21 Sept 2026Vals AI
LiveBench Agentic CodingAgentic64.6%77.525 Jun 2026LiveBench
PerceptionBench (Internal)Multimodal63.5%Moonshot AI
OpenHarmony Bench v1.0Coding60.8%70.0OpenHarmony Bench authors
ProofBench v1.1Math58.0%72.921 Sept 2026Vals AI
DeepSWEAgentic56.6%64.9Datacurve AI
Humanity's Last Exam with toolsAgentic56.2%69.0DeepSeek-AI
NL2RepoCoding55.9%68.9MiniMax
τ²-bench BankingAgentic55.2%38.4xhigh effort · Sierra4 Aug 2026Sierra Research
JobBenchAgentic53.4%70.8Yuetai Li et al.
Agents' Last ExamAgentic52.4%90.5DeepSeek-AI
ZeroBench_main with PythonMultimodal49.0%Moonshot AI / ZeroBench authors
FrontierMath-Tier-4-v2-PrivateMath46.3%70.2xhigh effortEpoch AI
SimpleQA VerifiedKnowledge45.8%64.0xhigh effortEpoch AI
Humanity's Last ExamKnowledge43.6%65.7Center for AI Safety et al.
Humanity's Last Exam without toolsKnowledge43.6%65.7OpenAI
SkillsBenchCoding42.0%58.6OpenHands11 Sept 2026Vals AI
MLS-Bench LiteCoding41.0%MLS-Bench
Mystery Game PuzzlesReasoning38.0%74.3xhigh effortEpoch AI
Chess PuzzlesReasoning29.0%61.0xhigh effortEpoch AI
AutomationBenchAgentic27.3%60.7Moonshot AI
Terminal-Bench 4.0Agentic24.7%76.921 Sept 2026Vals AI
ZeroBenchMultimodal24.0%Meta AI
Code MigrationCoding24.0%60.221 Sept 2026Vals AI
OSWorld 2.0Agentic19.4%65.2Mengqi Yuan et al.
FrontierSWE v2Coding15.8%63.2Proximal
Vibe Code Bench 1-100Coding12.8%65.9OpenHands16 Sept 2026Vals AI
ProgramBenchCoding0.0%21 Sept 2026Vals AI

48 benchmarks count, from 56 of 84 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsEpoch AI, collected directlyCC BY — free to use and redistribute with attributionVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsSierra Research, collected directlyMIT — results are in the licensed repository

More from Alibaba

Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6Qwen3.7 Plus58.2