MiMo-V2.6-Pro

MiMo-V2.6-Pro is a reasoning model from Xiaomi in the MiMo-V2.6 family. 19 benchmarks count toward its score, in 4 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.73.5 ±6.7
CoverageShare of the index weight with results.70%
SpeedOutput tokens per second.38/s
Input / 1MUS dollars per 1M input tokens.$0.435
Output / 1MUS dollars per 1M output tokens.$0.87
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
76.3
CodingCode writing and repair.
58.5
ReasoningLogic problems and puzzles.
81.0
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
71.8
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

19 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
CyberGymAgentic94.0%80.5Zhun Wang et al.
Terminal-Bench 2.1 (provider run)Agentic89.9%77.1DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic89.9%77.1DeepSeek-AI
Artificial Analysis Long Context ReasoningReasoning86.3%67.9Artificial Analysis
OSWorld-VerifiedAgentic82.0%70.1Tianbao Xie et al.
Toolathlon-VerifiedAgentic76.9%74.7Moonshot AI
DeepSWEAgentic71.9%76.0Datacurve AI
JobBenchAgentic62.0%76.7Yuetai Li et al.
Artificial Analysis SciCodeCoding60.9%77.4Artificial Analysis
GDPval-AA normalizedAgentic58.7%80.6Artificial Analysis
Artificial Analysis AutomationBenchAgentic58.6%70.1Artificial Analysis
AutomationBenchAgentic53.1%95.0Moonshot AI
Artificial Analysis Humanity's Last ExamKnowledge49.4%78.1Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge46.3%80.3Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge34.8%56.9Artificial Analysis
Agents' Last ExamAgentic31.6%71.0DeepSeek-AI
Critical Physics TasksReasoning26.6%94.2Artificial Analysis
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding26.5%39.7John Yang et al.
Artificial Analysis GDP.pdfAgentic19.2%70.3Artificial Analysis
ExploitGymAgentic17.8%74.0Zhun Wang et al.

19 benchmarks count, from 20 of 20 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence stated

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6

More from Xiaomi

MiMo-V2.5-Pro58.1MiMo-V2.555.7MiMo-V2-Pro53.6MiMo-V2-Omni49.1MiMo-V2-Flash43.9