Qwen3.6-27B

Qwen3.6-27B is a reasoning model from Alibaba. 48 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.53.8 ±2.3
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.22/s
Input / 1MUS dollars per 1M input tokens.$0.3
Output / 1MUS dollars per 1M output tokens.$2
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
58.5
CodingCode writing and repair.
54.9
ReasoningLogic problems and puzzles.
46.5
MultimodalTasks with images and text.
53.6
KnowledgeFacts and expert knowledge.
50.6
MultilingualTasks in many languages.
83.9
InstructionTasks with strict rules in the prompt.
51.0
MathMath problems.
52.2

Results

48 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
CountBenchMultimodal97.8%Qwen
V*Multimodal94.7%54.9Z.AI
τ²-Bench Tool-Agent-User EvaluationAgentic94.2%66.4Victor Barres et al.
AIME 2026Math94.1%53.5Qwen
Harvard-MIT Mathematics Tournament February 2025Math93.8%52.4Qwen
MMLU-ReduxKnowledge93.5%51.6Qwen
RefCOCO averageMultimodal92.5%RefCOCO dataset authors
C-EvalKnowledge91.4%C-Eval authors
Harvard-MIT Mathematics Tournament November 2025Math90.7%Qwen
Graduate-Level Google-Proof Q&AKnowledge87.8%59.3David Rein et al.
Video-MME with subtitleMultimodal87.7%Qwen
MLVU mean averageMultimodal86.6%Qwen
Massive Multitask Language Understanding ProfessionalKnowledge86.2%56.4Yubo Wang et al.
DynaMathMultimodal85.6%Qwen
GPQA diamondKnowledge84.8%56.6none effortEpoch AI
VideoMMMUMultimodal84.4%Qwen
Harvard-MIT Mathematics Tournament February 2026Math84.3%53.2Qwen
Artificial Analysis GPQA DiamondKnowledge84.2%54.8Artificial Analysis
RealWorldQAMultimodal84.1%54.5Qwen
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeCoding83.9%60.6Naman Jain et al.
Massive Multi-discipline Multimodal UnderstandingMultimodal82.9%52.1MMMU authors
MStarMultimodal81.4%Qwen
CC-OCRMultimodal81.2%Qwen
MMAnswerBenchMath80.8%Qwen
LiveBench MathematicsMath79.9%54.625 Jun 2026LiveBench
CharXiv ReasoningMultimodal78.4%55.3CharXiv authors
Artificial Analysis Long Context ReasoningReasoning77.3%61.7Artificial Analysis
Software Engineering Benchmark VerifiedCoding77.2%59.5Carlos E. Jimenez et al.
Massive Multi-discipline Multimodal Understanding ProMultimodal75.8%50.7MMMU-Pro authors
Artificial Analysis MMMU-ProMultimodal74.6%58.3Artificial Analysis
Claw-EvalAgentic72.4%72.2Bowen Ye et al.
LiveBench CodingCoding71.8%56.925 Jun 2026LiveBench
LiveBench Data AnalysisReasoning70.4%53.825 Jun 2026LiveBench
AndroidWorldAgentic70.3%Z.AI
LiveBench ReasoningReasoning70.3%53.625 Jun 2026LiveBench
SWE-benchCoding70.0%53.71 Sept 2026Vals AI
Artificial Analysis IFBenchInstruction67.6%57.7Artificial Analysis
OTIS Mock AIME 2024-2025Math66.7%48.8none effortEpoch AI
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge66.0%52.1Xiaoxuan Du et al.
EuroEval PolishMultilingual65.3%83.9EuroEval
LiveBench LanguageKnowledge63.3%47.025 Jun 2026LiveBench
ERQAMultimodal62.5%57.3Qwen
SimpleVQAMultimodal56.1%45.9Z.AI
Gert Labs Composite Game BenchmarkAgentic54.8%61.2Gert Labs
Artificial Analysis Coding IndexCoding53.7%56.8Artificial Analysis
SWE-bench ProCoding53.5%55.7Xiang Deng et al.
QwenClawBenchAgentic53.4%52.8Qwen
LiveBench Instruction FollowingInstruction53.2%44.325 Jun 2026LiveBench
Terminal-Bench 2.0Agentic44.9%54.84 Jun 2026Vals AI
Artificial Analysis SciCodeCoding42.8%52.3Artificial Analysis
LiveBench Agentic CodingAgentic39.3%53.725 Jun 2026LiveBench
NL2RepoCoding36.2%53.2MiniMax
FrontierMath-Tiers-1-3-v2-PrivateMath34.0%50.6none effortEpoch AI
Humanity's Last ExamKnowledge24.0%49.1Center for AI Safety et al.
GDPval-AA normalizedAgentic23.7%53.8Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge23.1%49.6Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge21.4%49.2Artificial Analysis
Artificial Analysis Agentic IndexAgentic20.1%53.5Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge19.6%38.1Artificial Analysis
Vibe Code Bench v1.1Coding11.9%47.1OpenHands21 Sept 2026Vals AI
Chess PuzzlesReasoning9.0%35.2none effortEpoch AI
Mystery Game PuzzlesReasoning7.0%41.4none effortEpoch AI
Critical Physics TasksReasoning1.1%40.7Artificial Analysis

48 benchmarks count, from 51 of 63 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionLiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AIEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repository

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.8-27B63.1Qwen3.7 Max62.6