DeepSeek
availableShows if the model has enough results for an index.DeepSeek V4 Flash 0731
DeepSeek V4 Flash 0731 is a reasoning model from DeepSeek in the DeepSeek V4 Flash family. 29 benchmarks count toward its score, in 5 categories.
IndexOverall score. 50 is the middle.64.6 ±5.0
CoverageShare of the index weight with results.80%
SpeedOutput tokens per second.221/s
Input / 1MUS dollars per 1M input tokens.$0.14
Output / 1MUS dollars per 1M output tokens.$0.28
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A
50 is the middle of the board. The range shows the doubt in the index.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
29 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| Harvard-MIT Mathematics Tournament February 2026 | Math | 94.8% | 61.1 | — | — | Qwen |
| LiveCodeBench Pass@1 with Chain-of-Thought | Coding | 91.6% | — | — | — | DeepSeek |
| Artificial Analysis GPQA Diamond | Knowledge | 90.8% | 61.6 | — | — | Artificial Analysis |
| VulcanBench v3 | Coding | 88.4% | 75.0 | — | — | VulcanBench contributors |
| IMOAnswerBench | Math | 88.4% | — | — | — | DeepSeek-AI |
| Graduate-Level Google-Proof Q&A | Knowledge | 88.1% | 59.6 | — | — | David Rein et al. |
| GPQA Diamond | Knowledge | 88.1% | 59.6 | — | — | David Rein et al. |
| Massive Multitask Language Understanding Professional | Knowledge | 86.2% | 56.4 | — | — | Yubo Wang et al. |
| Apex Shortlist | Math | 85.7% | — | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 82.7% | 72.9 | — | — | DeepSeek-AI |
| Terminal-Bench 2.1 (provider run) | Agentic | 82.7% | 72.9 | — | — | DeepSeek-AI |
| Artificial Analysis Long Context Reasoning | Reasoning | 79.7% | 63.4 | — | — | Artificial Analysis |
| Software Engineering Benchmark Verified | Coding | 79.0% | 60.9 | — | — | Carlos E. Jimenez et al. |
| Chinese-SimpleQA | Knowledge | 78.9% | — | — | — | DeepSeek-AI |
| MRCR 1M | Reasoning | 78.7% | — | — | — | DeepSeek-AI |
| CyberGym | Agentic | 76.7% | 68.7 | — | — | Zhun Wang et al. |
| BrowseComp | Agentic | 73.2% | 61.3 | — | — | OpenAI |
| Toolathlon-Verified | Agentic | 70.3% | 69.1 | — | — | Moonshot AI |
| Artificial Analysis Coding Index | Coding | 69.1% | 67.6 | — | — | Artificial Analysis |
| MCP Atlas | Agentic | 69.0% | 60.1 | — | — | OpenAI |
| DeepSeek DSBench FullStack | Coding | 68.7% | — | — | — | DeepSeek-AI |
| CorpusQA 1M | Reasoning | 60.5% | — | — | — | DeepSeek-AI |
| DeepSeek DSBench Hard | Coding | 59.6% | — | — | — | DeepSeek-AI |
| DeepSWE | Agentic | 54.4% | 63.3 | — | — | Datacurve AI |
| NL2Repo | Coding | 54.2% | 67.6 | — | — | MiniMax |
| OpenHarmony Bench v1.0 | Coding | 53.8% | 62.7 | — | — | OpenHarmony Bench authors |
| SWE-bench Pro | Coding | 52.6% | 54.9 | — | — | Xiang Deng et al. |
| Artificial Analysis SciCode | Coding | 50.3% | 62.7 | — | — | Artificial Analysis |
| Toolathlon | Agentic | 47.8% | 62.5 | — | — | OpenAI |
| GDPval-AA normalized | Agentic | 46.3% | 71.1 | — | — | Artificial Analysis |
| Humanity's Last Exam with tools | Agentic | 45.1% | 58.3 | — | — | DeepSeek-AI |
| Artificial Analysis Agentic Index | Agentic | 41.7% | 71.2 | — | — | Artificial Analysis |
| Artificial Analysis Omniscience Accuracy | Knowledge | 40.4% | 63.8 | — | — | Artificial Analysis |
| Artificial Analysis Humanity's Last Exam | Knowledge | 38.6% | 66.4 | — | — | Artificial Analysis |
| Humanity's Last Exam | Knowledge | 34.8% | 58.3 | — | — | Center for AI Safety et al. |
| Artificial Analysis Intelligence Index | Knowledge | 34.3% | 65.3 | — | — | Artificial Analysis |
| Measuring Short-Form Factuality in Large Language Models | Knowledge | 34.1% | — | — | — | Jason Wei et al. |
| Apex | Math | 33.0% | — | — | — | DeepSeek-AI |
| Agents' Last Exam | Agentic | 25.2% | 64.9 | — | — | DeepSeek-AI |
| AutomationBench | Agentic | 25.1% | 56.9 | — | — | Moonshot AI |
| Critical Physics Tasks | Reasoning | 16.6% | 73.2 | — | — | Artificial Analysis |
29 benchmarks count, from 31 of 41 results. A grey row does not count. Too few models took that benchmark.