Kimi K2.5 (Reasoning)

Kimi K2.5 (Reasoning) is a reasoning model from Moonshot AI in the Kimi K2.5 family. 19 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.54.6 ±4.2
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.$0.6
Output / 1MUS dollars per 1M output tokens.$3
ContextMaximum tokens in one request.128K
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
51.6
CodingCode writing and repair.
55.5
ReasoningLogic problems and puzzles.
53.5
MultimodalTasks with images and text.
57.2
KnowledgeFacts and expert knowledge.
57.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
60.3
MathMath problems.
52.2

Results

19 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
American Invitational Mathematics Examination 2025Math96.1%52.2Mathematical Association of America
τ²-Bench Tool-Agent-User EvaluationAgentic95.9%67.6Victor Barres et al.
Artificial Analysis GPQA DiamondKnowledge87.9%58.6Artificial Analysis
Graduate-Level Google-Proof Q&AKnowledge87.6%59.1David Rein et al.
Massive Multitask Language Understanding ProfessionalKnowledge87.1%57.8Yubo Wang et al.
Massive Multi-discipline Multimodal Understanding ProMultimodal78.5%55.1MMMU-Pro authors
Artificial Analysis Long Context ReasoningReasoning78.0%62.2Artificial Analysis
Software Engineering Benchmark VerifiedCoding76.8%59.1Carlos E. Jimenez et al.
Artificial Analysis MMMU-ProMultimodal75.4%59.2Artificial Analysis
Artificial Analysis IFBenchInstruction70.2%60.3Artificial Analysis
BrowseCompAgentic60.6%50.8OpenAI
Artificial Analysis Coding IndexCoding46.8%51.9Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge35.2%57.4Artificial Analysis
Gert Labs Composite Game BenchmarkAgentic32.6%41.5Gert Labs
Artificial Analysis Humanity's Last ExamKnowledge30.7%57.8Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge23.5%51.7Artificial Analysis
GDPval-AA normalizedAgentic16.2%48.0Artificial Analysis
APEX-Agents-AAAgentic11.5%50.2Artificial Analysis / Mercor
Critical Physics TasksReasoning3.1%44.9Artificial Analysis

19 benchmarks count, from 19 of 19 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Moonshot AI

Kimi K372.0Kimi K2.660.9Kimi K2.7 Code59.5Kimi K2.553.3Kimi K2.5 Thinking53.8