DeepSeek V4 Pro 0813

DeepSeek V4 Pro 0813 is a reasoning model from DeepSeek in the DeepSeek V4 Pro family. 32 benchmarks count toward its score, in 6 categories.

availableShows if the model has enough results for an index.
IndexOverall score. 50 is the middle.68.8 ±4.4
CoverageShare of the index weight with results.85%
SpeedOutput tokens per second.68/s
Input / 1MUS dollars per 1M input tokens.$0.435
Output / 1MUS dollars per 1M output tokens.$0.87
ContextMaximum tokens in one request.1M
EloLMArena rating and rank.N/A

50 is the middle of the board. The range shows the doubt in the index.

CapabilitiesScore per category. 50 is the middle.

50 is the middle
AgenticMulti-step tasks with tools.
70.3
CodingCode writing and repair.
65.4
ReasoningLogic problems and puzzles.
70.0
MultimodalTasks with images and text.
N/A
KnowledgeFacts and expert knowledge.
68.7
MultilingualTasks in many languages.
N/A
InstructionTasks with strict rules in the prompt.
66.7
MathMath problems.
61.4

Results

32 counted
BenchmarkThe test name.CategoryThe capability that the test measures.ResultThe score from the publisher.IndexThis result on the index scale.RunThe settings of the run.DateDate of the result.Published byThe source of the result.
τ²-Bench Tool-Agent-User EvaluationAgentic96.2%67.8Victor Barres et al.
Harvard-MIT Mathematics Tournament February 2026Math95.2%61.4Qwen
LiveCodeBench Pass@1 with Chain-of-ThoughtCoding93.5%DeepSeek
Artificial Analysis GPQA DiamondKnowledge92.8%63.7Artificial Analysis
Apex ShortlistMath90.2%DeepSeek-AI
Graduate-Level Google-Proof Q&AKnowledge90.1%61.4David Rein et al.
GPQA DiamondKnowledge90.1%61.4David Rein et al.
IMOAnswerBenchMath89.8%DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic87.9%75.9DeepSeek-AI
Terminal-Bench 2.1 (provider run)Agentic87.9%75.9DeepSeek-AI
Massive Multitask Language Understanding ProfessionalKnowledge87.5%58.4Yubo Wang et al.
Chinese-SimpleQAKnowledge84.4%DeepSeek-AI
MRCR 1MReasoning83.5%DeepSeek-AI
BrowseCompAgentic83.4%69.8OpenAI
CyberGymAgentic83.3%73.2Zhun Wang et al.
Software Engineering Benchmark VerifiedCoding80.6%62.2Carlos E. Jimenez et al.
Artificial Analysis Long Context ReasoningReasoning80.3%63.8Artificial Analysis
Artificial Analysis IFBenchInstruction76.5%66.7Artificial Analysis
Toolathlon-VerifiedAgentic74.1%72.3Moonshot AI
MCP AtlasAgentic73.6%63.6OpenAI
DeepSeek DSBench FullStackCoding71.1%DeepSeek-AI
Artificial Analysis Coding IndexCoding68.8%67.5Artificial Analysis
DeepSeek DSBench HardCoding67.2%DeepSeek-AI
DeepSWEAgentic62.7%69.4Datacurve AI
CorpusQA 1MReasoning62.0%DeepSeek-AI
NL2RepoCoding61.5%73.4MiniMax
Humanity's Last Exam with toolsAgentic60.0%72.7DeepSeek-AI
OpenHarmony Bench v1.0Coding59.0%68.1OpenHarmony Bench authors
Measuring Short-Form Factuality in Large Language ModelsKnowledge57.9%Jason Wei et al.
SWE-bench ProCoding55.4%57.6Xiang Deng et al.
GDPval-AA normalizedAgentic54.5%77.4Artificial Analysis
Artificial Analysis Intelligence IndexKnowledge53.2%88.9Artificial Analysis
ToolathlonAgentic51.8%66.3OpenAI
Artificial Analysis SciCodeCoding51.0%63.6Artificial Analysis
Artificial Analysis EnterpriseOps-GymAgentic49.6%75.2Artificial Analysis
Artificial Analysis Agentic IndexAgentic49.6%77.6Artificial Analysis
Artificial Analysis Omniscience AccuracyKnowledge49.1%74.6Artificial Analysis
Humanity's Last ExamKnowledge42.7%65.0Center for AI Safety et al.
Artificial Analysis Humanity's Last ExamKnowledge41.0%69.0Artificial Analysis
ApexMath38.3%DeepSeek-AI
AutomationBenchAgentic31.8%68.4Moonshot AI
Agents' Last ExamAgentic25.7%65.4DeepSeek-AI
APEX-Agents-AAAgentic24.3%60.3Artificial Analysis / Mercor
Critical Physics TasksReasoning18.0%76.1Artificial Analysis

32 benchmarks count, from 34 of 44 results. A grey row does not count. Too few models took that benchmark.

Sources

BenchLM benchmark aggregationUsed with attribution; per-benchmark results credited to their original authors

Same level, lower price

Qwen3.8-Flash-Next66.4 · FreeGPT-6 Luna68.4 · $0.5DeepSeek V4.1 Flash70.6 · $0.6

More from DeepSeek

DeepSeek V4.1 Flash70.6DeepSeek V4 Flash 073164.6DeepSeek V4 Pro 042362.4DeepSeek V4 Flash 042363.0DeepSeek V3.2 (Thinking)49.9