Qwen3.8-Flash-Next

Qwen3.8-Flash-Next is a reasoning model from Alibaba. 31 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.66.4 ±3.4
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.54/s
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
72.9
CodingCode writing and repair.
63.3
ReasoningLogic problems and puzzles.
64.4
MultimodalTasks with images and text.
65.9
KnowledgeFacts and expert knowledge.
61.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
71.1
MathMath problems.
62.5

Results

31 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MathVision with PythonMultimodal95.7%Moonshot AI / MathVision authors
Artificial Analysis GPQA DiamondKnowledge92.3%63.1Artificial Analysis
LiveCodeBench v6Coding91.9%60.7LiveCodeBench maintainers
Graduate-Level Google-Proof Q&AKnowledge91.7%62.9David Rein et al.
GPQA DiamondKnowledge91.7%62.9David Rein et al.
MathVisionMultimodal90.6%Qwen
CharXiv ReasoningMultimodal90.6%69.3CharXiv authors
RealWorldQAMultimodal88.5%63.2Qwen
LiveBench ReasoningReasoning87.4%77.325 Jun 2026LiveBench
LiveBench MathematicsMath85.8%62.525 Jun 2026LiveBench
CharXiv Reasoning without toolsMultimodal84.6%CharXiv authors
AndroidWorldAgentic84.5%Z.AI
Instruction Following BenchmarkInstruction81.3%60.6Benchmark authors
Artificial Analysis MMMU-ProMultimodal79.8%64.6Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning79.7%63.4Artificial Analysis
LiveBench Instruction FollowingInstruction77.1%81.625 Jun 2026LiveBench
LVBenchMultimodal76.6%Qwen Team
LiveBench LanguageKnowledge74.6%60.625 Jun 2026LiveBench
LiveBench Data AnalysisReasoning74.2%59.125 Jun 2026LiveBench
CoWorkBenchAgentic73.9%Qwen Team
Toolathlon-VerifiedAgentic73.5%71.8Moonshot AI
Artificial Analysis Coding IndexCoding73.0%70.4Artificial Analysis
LiveBench CodingCoding72.6%58.125 Jun 2026LiveBench
ERQAMultimodal72.3%66.5Qwen
Vision2WebMultimodal64.0%Z.AI
SWE-bench ProCoding62.5%64.4Xiang Deng et al.
LiveBench Agentic CodingAgentic61.6%74.725 Jun 2026LiveBench
DeepSWEAgentic58.7%66.4Datacurve AI
JobBenchAgentic55.7%72.4Yuetai Li et al.
GDPval-AA normalizedAgentic55.6%78.3Artificial Analysis
Agents' Last ExamAgentic51.2%89.4DeepSeek-AI
Artificial Analysis SciCodeCoding50.6%63.1Artificial Analysis
NL2RepoCoding48.1%62.7MiniMax
Artificial Analysis Intelligence IndexKnowledge39.8%72.2Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge38.0%65.8Artificial Analysis
Humanity's Last ExamKnowledge35.9%59.2Center for AI Safety et al.
Humanity's Last Exam without toolsKnowledge35.9%59.2OpenAI
Artificial Analysis Omniscience AccuracyKnowledge24.5%44.2Artificial Analysis
OSWorld 2.0Agentic19.4%65.2Mengqi Yuan et al.
Critical Physics TasksReasoning11.1%61.7Artificial Analysis
Agent Arena task outcomeAgentic8.977.415 Sept 2026LMArena
Agent Arena command recoveryAgentic-2.065.115 Sept 2026LMArena
Agent Arena steerabilityAgentic-4.762.115 Sept 2026LMArena

31 benchmarks count, from 36 of 43 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

More from Alibaba

Qwen3.8 Max68.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6Qwen3.7 Plus58.2