DeepSeek
availableShows if the model has enough results for an index.DeepSeek V3.2
DeepSeek V3.2 is a non-reasoning model from DeepSeek. 26 benchmarks count toward its score, in 6 categories.
IndexOverall score. 50 is the middle.44.4 ±4.7
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.33/s
Input / 1MUS dollars per 1M input tokens.$0.269
Output / 1MUS dollars per 1M output tokens.$0.4
ContextMaximum tokens in one request.164K
EloLMArena rating and rank.1425 (#96)
50 is the middle of the board. The range shows the doubt in the index.
46,458 votes. Elo shows what people prefer. It does not change the score.
CapabilitiesScore per category. 50 is the middle.
50 is the middleResults
26 counted| BenchmarkThe test name. | CategoryThe capability that the test measures. | ResultThe score from the publisher. | IndexThis result on the index scale. | RunThe settings of the run. | DateDate of the result. | Published byThe source of the result. |
|---|---|---|---|---|---|---|
| τ²-bench Telecom | Agentic | 96.2% | 67.8 | DeepSeek | 2 Mar 2026 | Sierra Research |
| MGSM | Multilingual | 90.9% | — | — | 9 Jan 2026 | Vals AI |
| MMLU Pro | Knowledge | 83.1% | 51.4 | — | 1 Sept 2026 | Vals AI |
| τ²-bench Retail | Agentic | 81.1% | 57.0 | DeepSeek | 30 Apr 2026 | Sierra Research |
| τ²-Bench Tool-Agent-User Evaluation | Agentic | 78.9% | 55.4 | — | — | Victor Barres et al. |
| GPQA diamond | Knowledge | 77.3% | 49.6 | — | — | Epoch AI |
| GPQA Diamond | Knowledge | 76.3% | 48.7 | — | 1 Sept 2026 | Vals AI |
| Artificial Analysis GPQA Diamond | Knowledge | 75.1% | 45.5 | — | — | Artificial Analysis |
| React Native Evals | Coding | 71.5% | 47.8 | — | — | Callstack |
| SWE-bench Verified | Coding | 70.0% | 53.7 | high effort · mini-SWE-agent | 1 Sept 2026 | SWE-bench team |
| LiveCodeBench | Coding | 69.9% | 47.8 | — | 1 Sept 2026 | Vals AI |
| OTIS Mock AIME 2024-2025 | Math | 68.4% | 49.8 | — | — | Epoch AI |
| AIME | Math | 64.8% | 43.9 | — | 16 Apr 2026 | Vals AI |
| τ²-bench Airline | Agentic | 63.8% | 44.6 | DeepSeek | 2 Mar 2026 | Sierra Research |
| SWE-Rebench | Coding | 60.9% | — | — | — | Nebius |
| SWE-bench Multilingual | Multilingual | 59.0% | — | mini-SWE-agent | 20 Feb 2026 | SWE-bench team |
| SWE-bench Multilingual | Coding | 59.0% | — | mini-SWE-agent | 2 Sept 2026 | SWE-bench team |
| ARC-AGI-1 (semi-private) | Reasoning | 57.0% | 54.1 | — | — | ARC Prize Foundation |
| Terminal-Bench 1.0 | Agentic | 50.0% | 53.9 | — | 12 Jan 2026 | Vals AI |
| Artificial Analysis IFBench | Instruction | 49.0% | 38.8 | — | — | Artificial Analysis |
| Artificial Analysis Long Context Reasoning | Reasoning | 45.7% | 39.9 | — | — | Artificial Analysis |
| Claw-Eval | Agentic | 40.2% | 21.7 | — | — | Bowen Ye et al. |
| Terminal-Bench 2.0 | Agentic | 34.8% | 47.5 | — | 4 Jun 2026 | Vals AI |
| Gert Labs Composite Game Benchmark | Agentic | 29.6% | 38.9 | — | — | Gert Labs |
| Artificial Analysis Omniscience Accuracy | Knowledge | 24.0% | 43.6 | — | — | Artificial Analysis |
| FrontierMath-2025-02-28-Private | Math | 22.1% | 51.2 | — | — | Epoch AI |
| VITA-Bench | Agentic | 18.5% | 38.3 | — | — | Meituan LongCat Team |
| Artificial Analysis Intelligence Index | Knowledge | 16.0% | 42.5 | — | — | Artificial Analysis |
| IOI v1 | Coding | 14.4% | 48.1 | — | 9 Aug 2026 | Vals AI |
| Artificial Analysis Humanity's Last Exam | Knowledge | 11.2% | 36.7 | — | — | Artificial Analysis |
| Chess Puzzles | Reasoning | 7.5% | 33.2 | — | — | Epoch AI |
| ARC-AGI-2 (semi-private) | Reasoning | 4.0% | 41.7 | — | — | ARC Prize Foundation |
| FrontierMath-Tier-4-2025-07-01-Private | Math | 2.1% | 45.7 | — | — | Epoch AI |
| Critical Physics Tasks | Reasoning | 0.9% | 40.3 | — | — | Artificial Analysis |
26 benchmarks count, from 30 of 34 results. A grey row does not count. Too few models took that benchmark.