dots3-note Preview

dots3-note Preview is a reasoning model from Dots Studio in the dots3-note family. 18 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score out of 100.65.5 ±5.3
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.512K
EloLMArena rating and rank.N/A

The index is a score out of 100. The ± range shows how much it can change. A free tier is available.

CapabilitiesScore per category, out of 100.

Out of 100
AgenticMulti-step tasks with tools.
66.9
CodingCode writing and repair.
62.0
ReasoningLogic problems and puzzles.
78.5
MultimodalTasks with images and text.
60.7
KnowledgeFacts and expert knowledge.
73.3
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
56.6
MathMath problems.
N/A

Results

18 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result as a score out of 100.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
Instruction-Following EvalInstruction93.9%53.6Jeffrey Zhou et al.
DeepSearchQAAgentic92.1%71.3Meta AI
LiveCodeBench v6Coding91.5%60.3LiveCodeBench maintainers
IMOAnswerBenchMath90.9%DeepSeek-AI
MathVisionMultimodal87.7%Qwen
VideoMMMUMultimodal86.8%Qwen
BrowseCompAgentic83.3%69.7OpenAI
CharXiv Reasoning without toolsMultimodal83.1%CharXiv authors
Instruction Following BenchmarkInstruction80.4%59.6Benchmark authors
Multimodal Multi-disciplinary Video UnderstandingMultimodal79.9%MMVU benchmark maintainers
Massive Multi-discipline Multimodal Understanding ProMultimodal79.1%56.1MMMU-Pro authors
WideResearchAgentic78.9%64.7Qwen
Software Engineering Benchmark VerifiedCoding78.4%60.4Carlos E. Jimenez et al.
ARC-AGI-2 (semi-private)Reasoning76.8%78.5max effortARC Prize Foundation
Terminal-Bench 2.1 (provider run)Agentic75.1%68.4DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic75.1%68.4DeepSeek-AI
Claw-EvalAgentic73.4%73.7Bowen Ye et al.
SimpleVQAMultimodal72.5%65.3Z.AI
SWE-bench ProCoding61.0%63.0Xiang Deng et al.
GDP.pdf mean criteria pass rate without toolsMultimodal60.7%Surge AI and Anthropic
Toolathlon-VerifiedAgentic55.6%56.7Moonshot AI
PerceptionBench (Internal)Multimodal53.4%Moonshot AI
Humanity's Last Exam with toolsAgentic52.6%65.5DeepSeek-AI
Humanity's Last ExamKnowledge52.6%73.3Center for AI Safety et al.
BabyVisionMultimodal50.0%Meta AI
NL2RepoCoding49.8%64.1MiniMax
APEX-AgentsAgentic30.8%65.1Moonshot AI / APEX-Agents benchmark authors
ZeroBenchMultimodal19.0%Meta AI

18 benchmarks count, from 19 of 28 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authorsARC Prize Foundation, collected directlyNo licence stated. Their terms ask for written permission before commercial use