MiMo-V2-Omni

MiMo-V2-Omni is a reasoning model from Xiaomi. 11 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.49.1 ±6.4
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.N/A
Output / 1MUS dollars per 1M output tokens.N/A
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.1422 (#101)

The index is a score out of 100. The ± range shows how much it can change.

19,414 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
46.9
CodingCode writing and repair.
57.5
ReasoningLogic problems and puzzles.
50.4
MultimodalTasks with images and text.
52.5
KnowledgeFacts and expert knowledge.
48.0
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
43.4
MathMath problems.
N/A

Results

11 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic91.2%64.3Victor Barres et al.
Artificial Analysis GPQA DiamondKnowledge82.8%53.4Artificial Analysis
Artificial Analysis Long Context ReasoningReasoning75.0%60.1Artificial Analysis
Software Engineering Benchmark VerifiedCoding74.8%57.5Carlos E. Jimenez et al.
Artificial Analysis MMMU-ProMultimodal69.9%52.5Artificial Analysis
Artificial Analysis IFBenchInstruction53.5%43.4Artificial Analysis
Claw-EvalAgentic45.2%29.5Bowen Ye et al.
Artificial Analysis Intelligence IndexKnowledge23.9%52.3Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge22.1%48.5Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge19.3%37.7Artificial Analysis
Critical Physics TasksReasoning1.1%40.7Artificial Analysis

11 benchmarks count, from 11 of 11 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

More from Xiaomi

MiMo-V2.6-Pro73.5MiMo-V2.5-Pro58.1MiMo-V2.555.7MiMo-V2-Pro53.6MiMo-V2-Flash43.9