Qwen3.6 Plus

Qwen3.6 Plus is a reasoning model from Alibaba. 56 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.56.1 ±2.7
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.36/s
Input / 1MUS dollars per 1M input tokens.$0.325
Output / 1MUS dollars per 1M output tokens.$1.95
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1437 (#74)

50 is the middle of the board. The range shows the doubt in the index.

45,319 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
55.9
CodingCode writing and repair.
57.9
ReasoningLogic problems and puzzles.
51.2
MultimodalTasks with images and text.
57.3
KnowledgeFacts and expert knowledge.
56.6
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
56.6
MathMath problems.
56.0

Results

56 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic97.7%68.9Victor Barres et al.
V*Multimodal96.9%57.6Z.AI
Harvard-MIT Mathematics Tournament February 2025Math96.7%54.9Qwen
AIME 2026Math95.3%54.4Qwen
Harvard-MIT Mathematics Tournament November 2025Math94.6%Qwen
AIMEMath94.6%57.816 Apr 2026Vals AI
MMLU-ReduxKnowledge94.5%53.2Qwen
Instruction-Following EvalInstruction94.3%54.4Jeffrey Zhou et al.
OTIS Mock AIME 2024-2025Math93.3%63.7Epoch AI
C-EvalKnowledge93.3%C-Eval authors
Graduate-Level Google-Proof Q&AKnowledge90.4%61.7David Rein et al.
Massive Multitask Language Understanding ProfessionalKnowledge88.5%60.0Yubo Wang et al.
GPQA diamondKnowledge88.4%59.8Epoch AI
Artificial Analysis GPQA DiamondKnowledge88.2%58.9Artificial Analysis
MathVisionMultimodal88.0%Qwen
Harvard-MIT Mathematics Tournament February 2026Math87.8%55.8Qwen
MMLU ProKnowledge87.7%58.71 Sept 2026Vals AI
GPQA DiamondKnowledge87.4%58.91 Sept 2026Vals AI
LiveCodeBench v6Coding87.1%56.3LiveCodeBench maintainers
Massive Multi-discipline Multimodal UnderstandingMultimodal86.0%55.1MMMU authors
LiveCodeBenchCoding86.0%62.51 Sept 2026Vals AI
MMLU-ProXMultilingual84.7%MMLU-ProX authors
MMMU ProMultimodal84.2%64.31 Sept 2026Vals AI
VideoMMMUMultimodal84.0%Qwen
MMAnswerBenchMath83.8%Qwen
LiveBench MathematicsMath83.7%59.725 Jun 2026LiveBench
CharXiv ReasoningMultimodal81.5%58.9CharXiv authors
Software Engineering Benchmark VerifiedCoding78.8%60.7Carlos E. Jimenez et al.
Massive Multi-discipline Multimodal Understanding ProMultimodal78.8%55.6MMMU-Pro authors
Artificial Analysis Long Context ReasoningReasoning78.3%62.4Artificial Analysis
LiveBench CodingCoding78.2%67.525 Jun 2026LiveBench
Artificial Analysis MMMU-ProMultimodal78.0%62.4Artificial Analysis
LiveBench ReasoningReasoning75.8%61.325 Jun 2026LiveBench
Instruction Following BenchmarkInstruction75.8%54.3Benchmark authors
Artificial Analysis IFBenchInstruction75.2%65.4Artificial Analysis
LiveBench LanguageKnowledge75.0%61.125 Jun 2026LiveBench
WideResearchAgentic74.3%59.7Qwen
MCP-TasksAgentic74.1%Qwen
SWE-benchCoding73.4%56.41 Sept 2026Vals AI
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge71.6%56.8Xiaoxuan Du et al.
τ³-Bench Tool-Agent-User EvaluationAgentic70.7%55.2Sierra Research
LiveBench Data AnalysisReasoning69.9%53.125 Jun 2026LiveBench
AI-NeedleReasoning68.3%Qwen
ScreenSpot ProMultimodal68.2%49.6Kaixin Li et al.
LongBench v2Reasoning62.0%LongBench v2 authors
Claw-EvalAgentic58.8%50.8Bowen Ye et al.
LiveBench Instruction FollowingInstruction58.3%52.325 Jun 2026LiveBench
NOVA-63Multilingual57.9%Qwen
SWE-Bench verifiedCoding57.9%43.9Epoch AI
QwenClawBenchAgentic57.2%56.2Qwen
SWE-bench ProCoding56.6%58.7Xiang Deng et al.
Artificial Analysis Coding IndexCoding54.5%57.4Artificial Analysis
Terminal-Bench 2.1Agentic53.2%55.421 Sept 2026Vals AI
Gert Labs Composite Game BenchmarkAgentic50.6%57.4Gert Labs
MCP AtlasAgentic48.2%44.6OpenAI
Terminal-Bench 2.0Agentic44.9%54.84 Jun 2026Vals AI
VITA-BenchAgentic44.3%59.7Meituan LongCat Team
SimpleQA VerifiedKnowledge44.1%62.4Epoch AI
DeepPlanningAgentic41.5%DeepPlanning authors
LiveBench Agentic CodingAgentic41.4%55.625 Jun 2026LiveBench
ToolathlonAgentic39.8%55.1OpenAI
FrontierMath-Tiers-1-3-v2-PrivateMath32.3%49.6none effortEpoch AI
Humanity's Last ExamKnowledge28.8%53.2Center for AI Safety et al.
Artificial Analysis Humanity's Last ExamKnowledge27.8%54.7Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge27.0%56.2Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge26.4%46.5Artificial Analysis
FrontierMath-2025-02-28-PrivateMath26.2%55.1Epoch AI
Vibe Code Bench v1.1Coding25.6%52.8OpenHands21 Sept 2026Vals AI
GDPval-AA normalizedAgentic23.8%53.8Artificial Analysis
ResearchClawBenchAgentic18.0%InternScience
Chess PuzzlesReasoning17.0%45.5Epoch AI
Mystery Game PuzzlesReasoning12.0%46.7none effortEpoch AI
Code MigrationCoding11.1%52.021 Sept 2026Vals AI
FrontierMath-Tier-4-2025-07-01-PrivateMath8.3%52.5Epoch AI
Critical Physics TasksReasoning2.9%44.5Artificial Analysis
ProgramBenchCoding0.0%21 Sept 2026Vals AI

56 benchmarks count, from 63 of 76 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not results

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6