Kimi K2.5

Kimi K2.5 is a non-reasoning model from Moonshot AI. 44 benchmarks count toward its score, in 7 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.53.3 ±2.9
CoverageShare of the index weight with results.95%
SpeedOutput tokens per second.47/s
Input / 1MUS dollars per 1M input tokens.$0.45
Output / 1MUS dollars per 1M output tokens.$2.25
ContextMaximum tokens in one request.262K
EloLMArena rating and rank.1446 (#53)

50 is the middle of the board. The range shows the doubt in the index.

70,513 votes. Elo shows what people prefer. It does not change the score.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
49.8
CodingCode writing and repair.
54.9
ReasoningLogic problems and puzzles.
50.0
MultimodalTasks with images and text.
57.2
KnowledgeFacts and expert knowledge.
56.1
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
57.0
MathMath problems.
54.8

Results

44 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
AIME25 first-party comparison snapshotMath96.3%Arcee AI
American Invitational Mathematics Examination 2025Math96.1%52.2Mathematical Association of America
τ²-Bench Tool-Agent-User EvaluationAgentic95.9%67.6Victor Barres et al.
AIME 2026Math95.8%54.8Qwen
Harvard-MIT Mathematics Tournament February 2025Math95.4%53.8Qwen
Instruction-Following EvalInstruction93.9%53.6Jeffrey Zhou et al.
OTIS Mock AIME 2024-2025Math92.2%63.1Epoch AI
Harvard-MIT Mathematics Tournament November 2025Math91.1%Qwen
Artificial Analysis GPQA DiamondKnowledge87.9%58.6Artificial Analysis
Graduate-Level Google-Proof Q&AKnowledge87.6%59.1David Rein et al.
GPQA DiamondKnowledge87.6%59.1David Rein et al.
GPQA diamondKnowledge87.6%59.1Epoch AI
Video-MMEMultimodal87.4%Video-MME benchmark team
Massive Multitask Language Understanding ProfessionalKnowledge87.1%57.8Yubo Wang et al.
MMLU-Pro first-party comparison snapshotKnowledge87.1%57.8Arcee AI
Harvard-MIT Mathematics Tournament February 2026Math87.1%55.3Qwen
VideoMMMUMultimodal86.6%Qwen
LiveCodeBench v6Coding85.0%54.4LiveCodeBench maintainers
MMLU-ProXMultilingual82.3%MMLU-ProX authors
MMAnswerBenchMath81.8%Qwen
Multimodal Multi-disciplinary Video UnderstandingMultimodal80.4%MMVU benchmark maintainers
Massive Multi-discipline Multimodal Understanding ProMultimodal78.5%55.1MMMU-Pro authors
Artificial Analysis Long Context ReasoningReasoning78.0%62.2Artificial Analysis
React Native EvalsCoding77.2%55.6Callstack
DeepSearchQAAgentic77.1%59.3Meta AI
Software Engineering Benchmark VerifiedCoding76.8%59.1Carlos E. Jimenez et al.
Artificial Analysis MMMU-ProMultimodal75.4%59.2Artificial Analysis
SWE-Bench verifiedCoding73.8%56.7Epoch AI
WideResearchAgentic72.7%57.9Qwen
SWE-bench Verified (mini-swe-agent-v2)Coding70.8%54.3Arcee AI
SWE-bench VerifiedCoding70.8%54.3high effort · mini-SWE-agent1 Sept 2026SWE-bench team
Artificial Analysis IFBenchInstruction70.2%60.3Artificial Analysis
SuperGPQA: Scaling LLM Evaluation Across 285 Graduate DisciplinesKnowledge69.2%54.8Xiaoxuan Du et al.
SWE-bench MultilingualMultilingual67.3%mini-SWE-agent20 Feb 2026SWE-bench team
SWE-bench MultilingualCoding67.3%mini-SWE-agent2 Sept 2026SWE-bench team
τ³-Bench Tool-Agent-User EvaluationAgentic65.7%51.0Sierra Research
ARC-AGI-1 (semi-private)Reasoning65.3%58.1ARC Prize Foundation
LongBench v2Reasoning61.0%LongBench v2 authors
BrowseCompAgentic60.6%50.8OpenAI
MCP-TasksAgentic59.1%Qwen
SWE-RebenchCoding58.5%Nebius
NOVA-63Multilingual56.0%Qwen
QwenClawBenchAgentic54.3%53.6Qwen
Claw-EvalAgentic52.3%40.7Bowen Ye et al.
SWE-bench ProCoding50.7%53.0Xiang Deng et al.
Scientific Code BenchmarkCoding48.7%59.2Benchmark authors
Artificial Analysis Coding IndexCoding46.8%51.9Artificial Analysis
Gert Labs Composite Game BenchmarkAgentic45.9%53.3Gert Labs
Artificial Analysis Omniscience AccuracyKnowledge35.2%57.4Artificial Analysis
SimpleQA VerifiedKnowledge34.3%53.3Epoch AI
Artificial Analysis Humanity's Last ExamKnowledge30.7%57.8Artificial Analysis
Humanity's Last ExamKnowledge30.1%54.3Center for AI Safety et al.
MCP AtlasAgentic29.5%30.5OpenAI
FrontierMath-2025-02-28-PrivateMath27.9%56.7Epoch AI
ToolathlonAgentic27.8%43.9OpenAI
Artificial Analysis Intelligence IndexKnowledge23.5%51.7Artificial Analysis
GDPval-AA normalizedAgentic16.2%48.0Artificial Analysis
DeepPlanningAgentic14.4%DeepPlanning authors
ResearchClawBenchAgentic14.0%InternScience
Chess PuzzlesReasoning12.0%39.0Epoch AI
ARC-AGI-2 (semi-private)Reasoning11.8%45.7ARC Prize Foundation
APEX-Agents-AAAgentic11.5%50.2Artificial Analysis / Mercor
JobBenchAgentic8.7%40.1Yuetai Li et al.
FrontierMath-Tier-4-2025-07-01-PrivateMath4.2%48.0Epoch AI
Critical Physics TasksReasoning3.1%44.9Artificial Analysis

44 benchmarks count, from 50 of 65 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsOpenRouter, collected directlyNo licence statedEpoch AI, collected directlyCC BY — free to use and redistribute with attributionSWE-bench team, collected directlyNo licence stated. The repository publishes submission records for reproducibility and transparency and asks that SWE-bench be citedARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use

Same level, lower price

Qwen3.8-Flash-Next66.4 · Freedots3-note Preview65.5 · FreeApodex 1.1 Mini59.1 · Free

More from Moonshot AI

Kimi K372.0Kimi K2.660.9Kimi K2.7 Code59.5Kimi K2.5 (Reasoning)54.6Kimi K2.5 Thinking53.8