Kimi K3

Kimi K3 is a reasoning model from Moonshot AI. 38 benchmarks count toward its score, in 5 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.72.0 ±4.6
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.37/s
Input / 1MUS dollars per 1M input tokens.$3
Output / 1MUS dollars per 1M output tokens.$15
ContextMaximum tokens in one request.1.05M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
72.6
CodingCode writing and repair.
68.4
ReasoningLogic problems and puzzles.
76.1
MultimodalTasks with images and text.
67.8
KnowledgeFacts and expert knowledge.
70.8
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
N/A
MathMath problems.
N/A

Results

38 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
MathVision with PythonMultimodal97.8%Moonshot AI / MathVision authors
DeepSearchQAAgentic95.0%73.6Meta AI
Artificial Analysis Harvey LAB-AAAgentic94.6%79.0Artificial Analysis
MathVisionMultimodal94.3%Qwen
Graduate-Level Google-Proof Q&AKnowledge93.5%64.6David Rein et al.
GPQA DiamondKnowledge93.5%64.6David Rein et al.
Artificial Analysis GPQA DiamondKnowledge93.5%64.4Artificial Analysis
CharXiv ReasoningMultimodal91.3%70.1CharXiv authors
BrowseCompAgentic91.2%76.3OpenAI
OmniDocBenchMultimodal91.1%Moonshot AI / OmniDocBench authors
Artificial Analysis Long Context ReasoningReasoning88.7%69.6Artificial Analysis
BabyVision with PythonMultimodal85.7%Moonshot AI
CharXiv Reasoning without toolsMultimodal84.8%CharXiv authors
MCP AtlasAgentic84.2%71.5OpenAI
MMMU-Pro with PythonMultimodal83.4%OpenAI
Massive Multi-discipline Multimodal Understanding ProMultimodal81.6%60.1MMMU-Pro authors
FrontierSWECoding81.2%Evan Chu et al.
Artificial Analysis MMMU-ProMultimodal80.5%65.5Artificial Analysis
ProgramBench: Can Language Models Rebuild Programs From Scratch?Coding77.8%75.3John Yang et al.
Artificial Analysis Coding IndexCoding76.2%72.7Artificial Analysis
VulcanBench v3Coding73.7%51.2VulcanBench contributors
DECK-Bench (Internal)Agentic73.5%Moonshot AI
Toolathlon-VerifiedAgentic73.2%71.5Moonshot AI
Kimi Code Bench v2Coding72.9%Moonshot AI
DeepSWEAgentic67.5%72.8Datacurve AI
OfficeQA ProMultimodal63.3%75.6OfficeQA Pro authors
cursorBench32Coding60.8%69.3Benchmark authors
Artificial Analysis SciCodeCoding59.5%75.4Artificial Analysis
PerceptionBench (Internal)Multimodal58.5%Moonshot AI
Artificial Analysis AutomationBenchAgentic58.3%69.8Artificial Analysis
OpenHarmony Bench v1.0Coding57.3%66.4OpenHarmony Bench authors
Humanity's Last ExamKnowledge56.0%76.2Center for AI Safety et al.
JobBenchAgentic52.9%70.5Yuetai Li et al.
GDPval-AA normalizedAgentic51.2%74.9Artificial Analysis
WorldVQA ForceAnswerMultimodal51.0%Moonshot AI / WorldVQA authors
Artificial Analysis Agentic IndexAgentic50.6%78.5Artificial Analysis
MLS-Bench LiteCoding48.3%MLS-Bench
Artificial Analysis ITBench-AAAgentic47.7%Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge47.6%72.8Artificial Analysis
Artificial Analysis Humanity's Last ExamKnowledge46.9%75.4Artificial Analysis
Artificial Analysis Tau3-BankingAgentic46.0%75.1Artificial Analysis
Artificial Analysis EnterpriseOps-GymAgentic45.3%69.4Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge43.6%76.9Artificial Analysis
Humanity's Last Exam without toolsKnowledge43.5%65.6OpenAI
SWE-MarathonCoding42.0%Abundant AI and BenchFlow
APEX-Agents-AAAgentic41.3%73.7Artificial Analysis / Mercor
ZeroBench_main with PythonMultimodal41.0%Moonshot AI / ZeroBench authors
Artificial Analysis AnalystAgentAgentic38.8%68.7Artificial Analysis
Medical Long Context Reasoning (MLCR-AA)Reasoning38.3%71.2Wisedocs and Artificial Analysis
APEX-AgentsAgentic37.6%70.7Moonshot AI / APEX-Agents benchmark authors
PostTrain BenchCoding36.6%Moonshot AI
SpreadsheetBench 2Agentic34.8%Moonshot AI
AutomationBenchAgentic30.8%66.7Moonshot AI
FrontierSWE v2Coding25.9%68.8Proximal
Critical Physics TasksReasoning23.4%87.4Artificial Analysis
ZeroBenchMultimodal23.0%Meta AI
Artificial Analysis GDP.pdfAgentic22.0%73.2Artificial Analysis
ApprenticeBench: end-to-end computer use, continual learning, and long-horizon agency on a real accounts-payable jobAgentic18.0%70.1NeoCognition

38 benchmarks count, from 40 of 58 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

DeepSeek V4.1 Flash70.6 · $0.6MiMo-V2.6-Pro73.5 · $0.87GPT-5.6 Luna69.9 · $1.2

More from Moonshot AI

Kimi K2.660.9Kimi K2.7 Code59.5Kimi K2.5 (Reasoning)54.6Kimi K2.553.3Kimi K2.5 Thinking53.8