Dots Studio
availableShows if the model has enough results for an index.dots3-note Preview
dots3-note Preview is a reasoning model from Dots Studio in the dots3-note family. 18 benchmarks count toward its score, in 6 categories.
IndexOverall score out of 100.65.5 ±5.3
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.—
Input / 1MUS dollars per 1M input tokens.Free
Output / 1MUS dollars per 1M output tokens.Free
ContextMaximum tokens in one request.512K
EloLMArena rating and rank.N/A
The index is a score out of 100. The ± range shows how much it can change. A free tier is available.
CapabilitiesScore per category, out of 100.
Out of 100Results
18 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result as a score out of 100. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| Instruction-Following Eval | Instruction | 93.9% | 53.6 | — | — | Jeffrey Zhou et al. |
| DeepSearchQA | Agentic | 92.1% | 71.3 | — | — | Meta AI |
| LiveCodeBench v6 | Coding | 91.5% | 60.3 | — | — | LiveCodeBench maintainers |
| IMOAnswerBench | Math | 90.9% | — | — | — | DeepSeek-AI |
| MathVision | Multimodal | 87.7% | — | — | — | Qwen |
| VideoMMMU | Multimodal | 86.8% | — | — | — | Qwen |
| BrowseComp | Agentic | 83.3% | 69.7 | — | — | OpenAI |
| CharXiv Reasoning without tools | Multimodal | 83.1% | — | — | — | CharXiv authors |
| Instruction Following Benchmark | Instruction | 80.4% | 59.6 | — | — | Benchmark authors |
| Multimodal Multi-disciplinary Video Understanding | Multimodal | 79.9% | — | — | — | MMVU benchmark maintainers |
| Massive Multi-discipline Multimodal Understanding Pro | Multimodal | 79.1% | 56.1 | — | — | MMMU-Pro authors |
| WideResearch | Agentic | 78.9% | 64.7 | — | — | Qwen |
| Software Engineering Benchmark Verified | Coding | 78.4% | 60.4 | — | — | Carlos E. Jimenez et al. |
| ARC-AGI-2 (semi-private) | Reasoning | 76.8% | 78.5 | max effort | — | ARC Prize Foundation |
| Terminal-Bench 2.1 (provider run) | Agentic | 75.1% | 68.4 | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 75.1% | 68.4 | — | — | DeepSeek-AI |
| Claw-Eval | Agentic | 73.4% | 73.7 | — | — | Bowen Ye et al. |
| SimpleVQA | Multimodal | 72.5% | 65.3 | — | — | Z.AI |
| SWE-bench Pro | Coding | 61.0% | 63.0 | — | — | Xiang Deng et al. |
| GDP.pdf mean criteria pass rate without tools | Multimodal | 60.7% | — | — | — | Surge AI and Anthropic |
| Toolathlon-Verified | Agentic | 55.6% | 56.7 | — | — | Moonshot AI |
| PerceptionBench (Internal) | Multimodal | 53.4% | — | — | — | Moonshot AI |
| Humanity's Last Exam with tools | Agentic | 52.6% | 65.5 | — | — | DeepSeek-AI |
| Humanity's Last Exam | Knowledge | 52.6% | 73.3 | — | — | Center for AI Safety et al. |
| BabyVision | Multimodal | 50.0% | — | — | — | Meta AI |
| NL2Repo | Coding | 49.8% | 64.1 | — | — | MiniMax |
| APEX-Agents | Agentic | 30.8% | 65.1 | — | — | Moonshot AI / APEX-Agents benchmark authors |
| ZeroBench | Multimodal | 19.0% | — | — | — | Meta AI |
18 benchmarks count, from 19 of 28 results. A grey row does not count. Too few models took that benchmark.