Qwen3.8-27B

Qwen3.8-27B is a reasoning model from Alibaba. 47 benchmarks count toward its score, in 8 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.63.1 ±2.4
CoverageShare of the index weight with results.100%
SpeedOutput tokens per second.44/s
Input / 1MUS dollars per 1M input tokens.$0.42
Output / 1MUS dollars per 1M output tokens.$3
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.1439 (#68)

50 is the middle of the board. The range shows the doubt in the index.

10,697 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
69.0
CodingCode writing and repair.
61.7
ReasoningLogic problems and puzzles.
59.0
MultimodalTasks with images and text.
62.3
KnowledgeFacts and expert knowledge.
56.2
MultilingualTasks in many languages.
83.6
InstructionTasks with strict rules in the prompt.
66.6
MathMath problems.
59.5

Results

47 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MathVision with PythonMultimodal94.6%Moonshot AI / MathVision authors
OmniDocBench 1.5Multimodal91.1%OpenAI
Artificial Analysis GPQA DiamondKnowledge90.5%61.3Artificial Analysis
LiveCodeBench v6Coding90.3%59.2LiveCodeBench maintainers
CharXiv ReasoningMultimodal90.2%68.8CharXiv authors
MathVisionMultimodal90.0%Qwen
Graduate-Level Google-Proof Q&AKnowledge89.2%60.6David Rein et al.
GPQA DiamondKnowledge89.2%60.6David Rein et al.
GPQA DiamondKnowledge88.9%60.31 Sept 2026Vals AI
LiveBench MathematicsMath86.2%63.125 Jun 2026LiveBench
SWE-benchCoding86.0%66.51 Sept 2026Vals AI
RealWorldQAMultimodal85.9%58.0Qwen
BabyVision with PythonMultimodal85.6%Moonshot AI
MMLU ProKnowledge84.3%53.41 Sept 2026Vals AI
OSWorld-VerifiedAgentic84.3%72.2Tianbao Xie et al.
LiveCodeBenchCoding84.0%60.71 Sept 2026Vals AI
MMMU ProMultimodal83.9%63.91 Sept 2026Vals AI
CharXiv Reasoning without toolsMultimodal83.7%CharXiv authors
VulcanBench v3Coding82.6%65.6VulcanBench contributors
Artificial Analysis Long Context ReasoningReasoning82.0%65.0Artificial Analysis
AndroidWorldAgentic81.9%Z.AI
LiveBench ReasoningReasoning80.0%67.125 Jun 2026LiveBench
Instruction Following BenchmarkInstruction79.5%58.5Benchmark authors
LiveBench Data AnalysisReasoning76.6%62.325 Jun 2026LiveBench
Artificial Analysis MMMU-ProMultimodal76.3%60.3Artificial Analysis
LiveBench CodingCoding75.7%63.325 Jun 2026LiveBench
LiveBench LanguageKnowledge74.3%60.325 Jun 2026LiveBench
Terminal-Bench 2.1 (provider run)Agentic73.0%67.1DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic73.0%67.1DeepSeek-AI
LiveBench Instruction FollowingInstruction72.7%74.725 Jun 2026LiveBench
CoWorkBenchAgentic70.7%Qwen Team
EuroEval PortugueseMultilingual68.5%87.8EuroEval
Artificial Analysis Coding IndexCoding68.1%66.9Artificial Analysis
EuroEval FrenchMultilingual66.8%85.7EuroEval
EuroEval ItalianMultilingual66.2%84.9EuroEval
BabyVisionMultimodal65.7%Meta AI
ERQAMultimodal65.5%60.1Qwen
EuroEval PolishMultilingual65.4%83.9EuroEval
EuroEval SwedishMultilingual64.9%83.3EuroEval
Vibe Code Bench v1.1Coding64.8%69.1OpenHands21 Sept 2026Vals AI
WebArena-Verified Browser Agent BenchmarkAgentic64.8%Amine El Hattami et al.
Vision2WebMultimodal62.9%Z.AI
EuroEval DutchMultilingual62.1%79.8EuroEval
SWE-bench ProCoding61.7%63.7Xiang Deng et al.
LiveBench Agentic CodingAgentic61.4%74.525 Jun 2026LiveBench
EuroEval SpanishMultilingual59.0%76.0EuroEval
Terminal-Bench 2.1Agentic58.4%58.521 Sept 2026Vals AI
EuroEval GermanMultilingual57.2%73.6EuroEval
Artificial Analysis Tau3-BankingAgentic48.0%77.8Artificial Analysis
Artificial Analysis SciCodeCoding46.6%57.5Artificial Analysis
Artificial Analysis Agentic IndexAgentic46.5%75.1Artificial Analysis
GDPval-AA normalizedAgentic45.4%70.4Artificial Analysis
Artificial Analysis EnterpriseOps-GymAgentic44.2%67.9Artificial Analysis
Agents' Last ExamAgentic42.9%81.6DeepSeek-AI
NL2RepoCoding42.3%58.1MiniMax
DeepSWEAgentic42.2%54.5Datacurve AI
IOICoding39.1%62.721 Sept 2026Vals AI
SkillsBenchCoding38.1%55.3OpenHands11 Sept 2026Vals AI
Artificial Analysis Humanity's Last ExamKnowledge33.9%61.3Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge33.7%64.5Artificial Analysis
JobBenchAgentic33.4%57.1Yuetai Li et al.
Humanity's Last ExamKnowledge30.8%54.9Center for AI Safety et al.
Humanity's Last Exam without toolsKnowledge30.8%54.9OpenAI
Medical Long Context Reasoning (MLCR-AA)Reasoning21.7%56.6Wisedocs and Artificial Analysis
ProofBench v1.1Math16.0%55.921 Sept 2026Vals AI
Artificial Analysis Omniscience AccuracyKnowledge15.6%33.2Artificial Analysis
Code MigrationCoding14.2%54.021 Sept 2026Vals AI
Critical Physics TasksReasoning5.4%49.7Artificial Analysis
Agent Arena task outcomeAgentic4.372.215 Sept 2026LMArena
Terminal-Bench 4.0Agentic4.0%62.321 Sept 2026Vals AI
Agent Arena steerabilityAgentic0.167.515 Sept 2026LMArena
ProgramBenchCoding0.0%21 Sept 2026Vals AI
Agent Arena command recoveryAgentic-5.261.515 Sept 2026LMArena

47 benchmarks count, from 62 of 73 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedVals AI, collected directlyNo licence stated. Read from the public leaderboard and credited to Vals AILiveBench, collected directlyNo licence stated for the leaderboard. The site repo serving these CSVs has no LICENSE; the harness repo carries upstream Apache-2.0 and MIT copies that cover code, not resultsEuroEval, collected directlyMIT — the leaderboard site and its CSV routes are in the licensed repositoryLMArena, collected directlyCC BY 4.0 (lmarena-ai/leaderboard-dataset on Hugging Face)

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeDeepSeek V4 Flash 042363.0 · $0.177

More from Alibaba

Qwen3.8 Max68.4Qwen3.8-Flash-Next66.4Qwen3.8 Max Preview69.0Qwen3.7 Max62.6Qwen3.7 Plus58.2